Datadory notebook

The Movies & Entertainment Data Guide (2026)

Datadory delivers movies & entertainment data covering every layer the industry gets asked about: 12 million-plus screen titles and 16 million-plus people joined on stable identifiers, a 32-million-rating recommender benchmark spanning January 1995 through October 2023, theatrical grosses reaching the early 1980s across domestic, international and worldwide views, 17,312 critic-ranked films carrying the review distributions behind their scores, and streaming offers resolved per title, per country and per provider across roughly 100 country editions - keyed, typed and delivered daily, weekly, or hourly, your call.

1,744 datasets. Pick your catch.

The movies & entertainment data landscape in 2026

This is the rare consumer industry that keeps receipts on itself. Every ticket becomes a gross, every review becomes a score, every subscription decision becomes a provider offer - and the result is measurement depth most industries never get near. Datadory catalogs 15 pooled datasets here, 14 primary records plus one adjacent television entry, scoring 10 down to 6 on the quality rubric with a mean of 7.79 against a 7.81 average across all 1,744 datasets in the catalog.

Ratings and recommender corpora form the second family. GroupLens MovieLens Datasets hold the 32-million-rating benchmark the field calibrates against, and the two Kaggle cuts - The Movies Dataset (TMDB + MovieLens) at 45,000 films joined to 26 million rating events, and the compact TMDB 5000 cut - arrive pre-linked so the modeling join is done before you open the file.

Box-office revenue splits by geography and grain: Box Office Mojo runs daily-to-yearly charts across domestic, international and worldwide views reaching the early 1980s, The Numbers tracks the domestic ledger with franchise roll-ups and distributor market share, and BFI Industry Data & Insights supplies the United Kingdom's official weekly Top 15 in pounds sterling.

Critic sentiment and streaming availability close the loop. Rotten Tomatoes Movies & TV and Metacritic Movie Browse put professional and audience verdicts beside the review counts that produced them; JustWatch Streaming Guide resolves where every title streams, rents, sells or plays ad-supported across roughly 100 country editions. A related television record, the TVmaze API - TV Show and Episode Database, extends the same identifier discipline into episodic TV.

Every record behind this guide documents its own field dictionary, sample rows and coverage span, so the grain is visible before anything lands in a model. Browse the full pool on the movies-entertainment data hub.

The core movies & entertainment datasets to know

Fourteen primary records sort into five working groups, and knowing which group owns a question decides which one you pull.

Title metadata

Ratings and recommender corpora

  • GroupLens MovieLens Datasets (quality score 10). The benchmark, and the reason this slice scores so well: 32,000,204 ratings from 200,948 anonymized users across 87,585 movies, spanning January 1995 through October 2023, with 2,000,072 free-text tag applications and a 10.5-million-score tag genome grading how well each label fits each film. Stable releases from 100K ratings to ml-32m, eleven fields documented end to end, and IMDb and TMDB crosswalk columns joined onto every movie row.
  • Kaggle - The Movies Dataset (TMDB + MovieLens) (quality score 8). The modeling join already done: 45,000 films of TMDB metadata - budgets, worldwide revenue, genres, cast, crew, plot keywords - pre-linked by identifier crosswalks to 26 million rating events from 270,000 users, across seven CSVs totalling roughly 944 MB uncompressed. Content features and collaborative signals arrive on shared movie IDs. Coverage freezes at its July 2017 upload, which makes it a reproducible benchmark rather than a live ledger.
  • Kaggle - TMDB 5000 Movie Dataset (quality score 6). The compact cut: roughly 4,800 films across two files joined on TMDB id, carrying production budgets, worldwide revenue, genres, keywords, production companies and countries, spoken languages, runtimes, TMDB ratings and vote counts, plus cast and crew credits with billing order and job roles.

Box-office revenue

  • Box Office Mojo (quality score 8). Longitudinal depth: fifteen documented fields on daily, weekend, weekly, monthly and yearly charts across domestic, international and worldwide views, with full gross histories for more than 20,000 releases and major-release dailies reaching the early 1980s. Theater counts, per-theater averages and distributors sit on every row, so trajectory curves are a filter rather than a reconstruction.
  • The Numbers - Movie Financial Data (quality score 7). The domestic ledger in motion: daily, weekend, weekly and yearly charts carrying rank, prior rank, gross, daily and weekly change, theater counts, per-theater averages, cumulative totals and days in release, plus franchise roll-ups and distributor market share. Figures are estimates by design; treat the headline as a range and the trend as the signal.
  • BFI Industry Data & Insights (quality score 8). Official statistics, citable as such: a weekly Top 15 weekend chart in pounds sterling with Friday-to-Sunday grosses reaching back to 2017, ten documented fields per row, one row per film per reporting week, and an annual Statistical Yearbook with editions since 2002. The only record here that anchors a UK screen-sector claim to a national statistical body.

Critic sentiment

  • Rotten Tomatoes Movies & TV (quality score 8). The only entry where professional and audience verdicts sit side by side per title: Tomatometer percentages from approved critics against Popcornmeter scores from verified ticket-buyers, with the counts behind both, Certified Fresh flags, advisory ratings, genres and cast, at one row per title and per season for TV.
  • Metacritic Movie Browse (quality score 7). Consensus shape rather than just a headline: 17,312 ranked films, each with the weighted Metascore, the 0-10 user score and critic reviews split into positive, neutral and negative counts with a derived sentiment label, plus release date, advisory rating, genres, runtime and artwork references. Classics from the 1920s sit beside 2026 releases.

Streaming availability

  • JustWatch Streaming Guide (quality score 7). The windowing layer, and the only record in the pool that answers which provider holds a title in which country: one normalized row per title, country and provider offer across roughly 100 country editions and several hundred services. Monetization type separates subscription, rental, purchase and ad-supported airings, and the scores viewers shop by ride the same row - IMDb rating and votes, TMDB popularity, the Tomatometer and JustWatch's own popularity score.

Television adjacency

  • TVmaze API - TV Show and Episode Database (related record). Episodic TV under the same identifier discipline: hundreds of thousands of shows worldwide with 27 verified fields each, per-day broadcast schedules filterable by country, complete season and episode spines under every show, cast and crew credits, and IMDb, TheTVDB and TVRage identifiers riding the same row.

Public-sector slice

  • Data.gov - Movies Tagged Datasets (quality score 6). The US federal catalog filtered to its movies tag: DCAT-US metadata records for government-published film datasets such as Film Locations in San Francisco and the Chicago Park District's Movies in the Parks series for 2014 through 2019, drawn from a catalog exceeding 552,000 datasets. Narrow, municipal, and occasionally exactly what a location-analysis project needs.

Read together, the fourteen cover a complete research arc - identity, audience, money, verdict and window - and the television neighbor extends the identity layer into episodic programming.

Movies & entertainment data by use case

Ten recurring jobs define how teams put this pool to work, and each maps to named records:

  1. Train recommender systems. Benchmark collaborative filtering on GroupLens MovieLens' 32,000,204 rating events and 10.5-million-score tag genome, then blend content features in from the Kaggle Movies Dataset's pre-linked TMDB credits and keywords.
  2. Build title search and enrichment. Anchor on the IMDb tconst spine, hang TMDB artwork, translations and watch-provider mappings off the crosswalk, and fall back to OMDb when one call needs the whole consolidated record - ratings array included.
  3. Compare UK and US theatrical performance. Put BFI's weekly Top 15 in pounds sterling beside Box Office Mojo's domestic weekend charts; both carry cinema or theater counts, so per-screen comparisons survive market-size differences.
  4. Study franchise and distributor economics. Mine The Numbers' franchise roll-ups and distributor market share, cross-checked against the distributor field on every Box Office Mojo row.
  5. Track streaming windows and exclusivity. JustWatch offer records separate subscription, rental, purchase and ad-supported airings per country, which turns exclusivity from anecdote into a queryable state.
  6. Test whether critics predict revenue. Pair Rotten Tomatoes' Tomatometer and Metacritic's Metascore distributions - 17,312 ranked films with review-count splits - against per-title gross histories once both sides sit on normalized titles.
  7. Analyse SVOD content strategy. Cut the Kaggle Netflix snapshot's 8,807 titles by genre label, production country and date added, holding the 2021 vintage as the fixed frame.
  8. Size the UK screen sector. Build on BFI Statistical Yearbook editions since 2002 - production, certification, audience and economy statistics - with the weekly chart supplying current-quarter texture.
  9. Locate filming-location records. Data.gov's movies-tagged slice surfaces municipal datasets such as SF Film Locations and Chicago's Movies in the Parks, ready for mapping work.
  10. Engineer alternative quant signals. Turn weekend-gross trajectories, theater counts and per-theater averages into features for wider models, treating headline figures as estimates and letting the trend carry the weight.

What separates a usable movies dataset from a raw dump

The strongest records in this pool share three traits, and each decides whether a project survives contact with production:

  • Identifier hygiene: the tconst resolves a title across ratings, alternate titles, crew, principals and episodes; nconst does the same for people; TMDB ids and IMDb crosswalks bridge the catalogs to each other. Joins stay clean when identifiers stay stable - which is why every record here leads with its key structure.
  • Depth of history: Box Office Mojo's charts reach the early 1980s, BFI's weekly archive reaches 2017 with a Yearbook line back to 2002, and MovieLens' ratings run January 1995 through October 2023 - long enough to fit models on full release cycles rather than fragments.
  • Declared caveats: The Numbers states its box-office figures are estimates subject to revision, the Kaggle bundles freeze at their upload dates, and the Netflix snapshot holds a fixed 2021 vintage. Name the caveats before you build on the numbers; they are part of the schema, not footnotes.

None of these traits shows up in a screenshot. They show up in sample rows - which is why every Datadory record leads with real rows and a documented field dictionary before any commitment.

Who uses movies & entertainment data?

Data scientists and ML engineers train recommenders and rankers on the MovieLens benchmark, calibrating content features from TMDB credits and keywords against 32 million collaborative signals. Their working rule - freeze a release before publishing results, because a moving benchmark proves nothing twice. Their page: data scientists.

Market researchers and consultants size markets and anchor claims to citable institutions: BFI official statistics for the UK chapter, box-office trackers for competitive benchmarks, critic panels for sentiment context. Their page: market researchers.

Investors and quant researchers build alternative signals - opening-weekend trajectories, theater-count expansion curves, per-theater averages - joined against studio and distributor fundamentals. Their rule of thumb: week-one figures for wide releases are noisy until sampling stabilizes, and estimated grosses deserve error bars, not decimals.

Competitive intelligence and product teams watch windowing moves - which titles go where, on what terms, for how long - through per-country offer records, and benchmark catalog strategy against the frozen SVOD snapshot.

Journalists, academics and students reach for the citable layer: national statistics produced by the British Film Institute, a ratings benchmark with a formal citation trail, and archives deep enough that a claim survives an editor.

Delivery: how the movies shelf reaches your warehouse

Two translation jobs belong on our side of the line rather than yours. First, entertainment sources arrive built for audiences, not analysts - turning browse surfaces and per-page records into typed columns with stable join keys is work done once upstream so you never repeat it. Second, the estimate-shaped records need interpretation attached: gross figures kept as estimates, review-derived ratios computed from raw counts, frozen snapshots labeled with their vintage. Both translations hold across every delivery.

What the numbers say

Every figure below comes straight from the cataloged records, as of August 2026. Against the 7.81 mean across all 1,744 datasets Datadory catalogs, this slice averages 7.79 with eight of fourteen primary records at 8 or higher - unusually strong, because the industry measured itself first. Set the pool against its alternatives in the best movies-entertainment datasets ranking.

Keep reading

Continue into the connected pages:

  1. movies-entertainment data hub - the full 15-record pool on one index page, with quality scores, field dictionaries, sample rows and coverage spans for every dataset in this guide.
  2. best movies-entertainment datasets - the ranked cut of this slice, with editorial tiebreaks and a scorecard of all ten leaders.
  3. data scientists - how modeling teams turn ratings corpora, tag vocabularies and credit graphs into recommenders and ranking systems.
Key numbers: Movies & Entertainment data in Datadory's catalog (as of August 2026)
measurefigure
Datasets cataloged for the industry15 pooled (14 primary + 1 related)
Mean quality score, industry slice7.79 of 10 (catalog-wide average: 7.81 across 1,744 datasets)
Records scoring 8 or higher8 of 14 primary records
Benchmark ratings panelMovieLens ml-32m: 32,000,204 ratings, Jan 1995 - Oct 2023, plus a 10.5-million-score tag genome
Longest gross historyBox Office Mojo: full histories for 20,000+ titles, dailies from the early 1980s
Ranked critic recordsMetacritic Movie Browse: 17,312 films with Metascores and review-count splits
Streaming markets coveredJustWatch Streaming Guide: roughly 100 country editions, several hundred providers
UK official statistics depthBFI weekly Top 15 from 2017; Statistical Yearbook editions since 2002
The five families at a glance: what each group of records answers
FamilyRecordsQuestion it answers
Ratings and recommender corporaGroupLens MovieLens Datasets, Kaggle - The Movies Dataset (TMDB + MovieLens), Kaggle - TMDB 5000 Movie DatasetWho watched, rated and tagged what - and will they like the next one?
Box-office revenueBox Office Mojo, The Numbers - Movie Financial Data, BFI Industry Data & InsightsHow much did it earn, where, on how many screens?
Critic sentimentRotten Tomatoes Movies & TV, Metacritic Movie BrowseWhat did critics and audiences actually say, and by how many reviews?
Streaming availabilityJustWatch Streaming Guide (plus TVmaze API as a related TV record)Where does it stream right now, on what terms, in which country?

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

How far back does box office history reach?

To the early 1980s for daily major-release data. Box Office Mojo carries full gross histories for more than 20,000 titles across daily, weekend, weekly, monthly and yearly views, each row holding rank, gross, theater count, per-theater average and distributor. The Numbers adds daily domestic ledgers with prior rank, days in release and cumulative totals, plus franchise roll-ups and distributor market share.

What data can power a movie recommendation model?

GroupLens MovieLens Datasets remain the benchmark: ml-32m holds 32,000,204 ratings from 200,948 anonymized users across 87,585 movies, spanning January 1995 through October 2023, beside 2,000,072 free-text tag applications and a 10.5-million-score tag genome. Kaggle - The Movies Dataset (TMDB + MovieLens) arrives with the modeling join already done - 45,000 films of budget, genre and credit metadata pre-linked to 26 million rating events.

Can I see where a title streams, country by country?

Through JustWatch Streaming Guide, the only windowing record in the pool: one normalized row per title, country and provider offer covers roughly 100 country editions and several hundred services, with monetization type distinguishing subscription, rental, purchase and ad-supported airings. The scores viewers actually shop by ride the same row - IMDb rating and vote count, TMDB popularity, the Tomatometer and JustWatch's own popularity score.

Do critic scores come with the review counts behind them?

Yes, on both sides of the aisle. Rotten Tomatoes Movies & TV pairs the Tomatometer and Popcornmeter with the critic and verified-audience volumes behind each percentage, plus Certified Fresh flags, at one row per title and per season for TV. Metacritic Movie Browse goes further on shape: 17,312 ranked films carry the weighted Metascore, the 0-10 user score and critic reviews split into positive, neutral and negative counts.

Which record covers UK film industry statistics?

BFI Industry Data & Insights, the British Film Institute's official read of the UK screen sector: a weekly Top 15 weekend chart in pounds sterling reaching back to 2017, ten documented fields per row - Rank, Film, Country of Origin, Weekend Gross, Distributor, week-on-week change, Weeks on release, Number of cinemas, Site average and Total Gross to date - and a Statistical Yearbook published annually since 2002.