Datadory notebook

Historical box office data download: four decades of grosses, delivered as rows

1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.

1,744 datasets. Pick your catch.

Which record carries the deepest US gross history?

Domestic depth is Box Office Mojo's job, and it is not close. Charts cut per day, weekend, week, month, quarter and year; domestic, per-territory international and worldwide views; all-time lifetime rankings; a release calendar covering upcoming wide, limited and re-release dates; and release groups that tie a title's related cuts into one lineage. Underneath the charts, per-release pages hold the time series itself - theater counts, per-theater averages, cumulative totals, days or weeks in release and the distributing studio on every observation - for roughly 20,000-plus titles, with daily figures reaching the early 1980s for major releases. It scores 8/10 on Datadory's rubric.

One delivered row reads like this:

```text # delivered grain: one row per release per chart period

Three facts arrive pre-decided in that shape. The gross lands as dollars with its period explicit, so "seventy million over a weekend" is what the cell says without unit math or calendar reconstruction. Saturation is priced natively - per_theater_avg lets a 4,500-screen tentpole and a 600-screen platform debut share one honest axis. And the cumulative column rides the same row as the weekly one, which turns a stack of weekly rows into a decay curve without diffing two pulls.

How far back do the daily archives reach?

A verified daily row, Wednesday, August 19, 2026:

What does the official UK layer add that US charts cannot?

Official statistics behave differently from tracker estimates, and BFI Industry Data & Insights is where that difference becomes useful. The record publishes film-level weekly UK box office as Top 15 weekend charts - roughly 52 releases a year, running 2017 to the present - produced as scheduled official statistics under the Statistics and Registration Services Act 2007, with an annual Statistical Yearbook whose editions reach back to 2002. Ten fields define the weekly row, and two of them exist nowhere in the US ledgers: country of origin, which turns UK-versus-US qualification into a filter rather than a research project, and site average in sterling, the exhibition-side read on how widely a title actually travelled.

Can budgets and revenues arrive in one bundle instead?

When the question is modelling rather than monitoring, the static bundle beats the chart walk - and saying so is not heresy. Kaggle - The Movies Dataset (TMDB + MovieLens) joins TMDB budget, revenue, runtime, genres, cast, crew and keywords for 45,000 films to 26 million MovieLens ratings from 270,000 users, across seven files with one row per movie in the metadata tables. Kaggle - TMDB 5000 Movie Dataset is the lighter variant: 4,803 movies, two files, roughly 46 MB uncompressed.

A row off the bundle shows why analysts keep reaching for it:

id             : 862
title          : Toy Story
release_date   : 1995-11-22
budget_usd     : 30,000,000
revenue_usd    : 373,554,033
runtime_min    : 81
vote_average   : 7.7    # 5,415 votes

Budget and revenue on one line is the return-on-production question answered without a join. The trade-off is the freeze, stated plainly rather than discovered late: the bundle covers films released on or before July 2017 and was assembled that November, so nothing newer appears in it and nothing in it revises afterwards. Datadory labels the vintage on every delivery and pairs the snapshot with the live ledgers above - history from the bundle, freshness from the feeds, both keyed to the same title identifiers. One convention to know before modelling: the economics files store cast, crew, genres and keywords as JSON arrays inside CSV columns, and deliveries normalise those into joinable child tables rather than leaving you regexing strings.

How do you pair grosses with critic scores?

Testing whether consensus predicts revenue takes one more join, and the slice supplies both sides of it. Rotten Tomatoes Movies & TV carries Tomatometer and audience scores with review counts across tens of thousands of title pages; Metacritic Movie Browse exposes 17,312 ranked movie records with weighted critic Metascores, user scores and review tallies. Join either to the gross ledgers - or to the bundle's revenue column - and the critic-versus-box-office panel assembles itself.

The join is keyed, not fuzzy. GroupLens MovieLens Datasets ships a links table mapping its movie identifiers to both major catalogs' IDs, so enrichment runs on equality joins instead of title-string guesswork. One coverage fact keeps such panels honest: the widest title universe in the pool - the 12-million-title IMDb-keyed corpus - carries identities, credits, ratings and alternate titles but no revenue field at all. Metadata breadth and gross history are different records here; use the corpus for who and what, the ledgers for how much.

Who builds on box office history, and how does delivery work?

Five jobs this history settles outright:

  1. Distribution and exhibition strategy - read holdover strength in the change and per-theater columns before committing screens next frame.
  2. Slate competitive tracking - rival openings land on the same charts within the cycle, so share shifts surface in days rather than at quarterly reports.
  3. Franchise valuation - release groups and series roll-ups turn sequels into comparable curves instead of anecdotes.
  4. Equity research on media companies - distributor market-share tables anchor theatrical revenue lines between filings.
  5. Citation-grade journalism and scholarship - any quoted figure resolves to title, chart and period, with the estimate status attached.

Where to go next

  1. Movies & entertainment data guide - the pillar walking all fifteen records in the slice, scored and compared.
  2. BFI versus the Kaggle bundle - official weekly statistics against a frozen analysis bundle, head to head.
  3. Best movies & entertainment datasets - the ranked view of the pool, and the movies & entertainment data hub for browsing every record's coverage, fields and delivery options.

When you're ready to build, request a sample cut to the titles, markets and date range on your desk this quarter. It arrives with the field dictionary attached - and the schema in the sample is the schema you ship against.

One verified UK row, decoded - weekly Top 15 chart, BFI Industry Data & Insights
FieldValueReading
FilmSpider-Man: Brand New DayTitle as released in UK cinemas
Country of OriginUK/USAQualification filter no US chart carries
Weekend GrossGBP 6,148,221Friday-to-Sunday gross in GBP, previews included where applicable
DistributorSony PicturesUK releasing studio - the join key for market-share work
Weeks on release3Run length; matches the US ledger's weeks-in-release column
Number of cinemas745Footprint for the weekend
Site averageGBP 8,253Weekend gross divided by cinemas - the exhibition-depth read
Total Gross to datePS78,755,985Cumulative UK gross since release
Which grain answers which question
GrainWhat it answersBest for
DailyVelocity and decay - what a title earned on a given day and how the curve bentOpening-trajectory models, holdover reads, event-driven research
Weekend and weeklyFrame-level competition - who won the weekend and by how muchSlate tracking, market-share narratives, press citations
Monthly, quarterly, yearlyAggregate demand without weekday noiseExecutive reporting, year-over-year comps, long-run history sweeps
Lifetime and all-timePecking order - where a title sits against every releaseFranchise valuation, catalog benchmarks, claim verification
Static bundleMoney beside attributes - budget against revenue on one rowModelling, ROI studies and teaching corpora; pair with live feeds for anything after 2017

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Movies & Entertainment United States/domestic market, per-territory international…

Box Office Mojo

rank · release_title · prior_period_rank …+8 more

Movies & Entertainment United States/domestic market as primary scope, with…

The Numbers - Movie Financial Data

rank · prev_rank · title …+8 more

Movies & Entertainment United Kingdom - UK box office and UK screen-sector…

BFI Industry Data & Insights

Rank · Film · Distributor

Movies & Entertainment Global films present in the GroupLens Full MovieLens dataset…

Kaggle - The Movies Dataset (TMDB + MovieLens)

imdb_id · movieId · imdbId …+13 more

Movies & Entertainment Global films listed on TMDB, predominantly English-language…

Kaggle - TMDB 5000 Movie Dataset

title · original_title · budget …+18 more

Movies & Entertainment Global rater base

GroupLens MovieLens Datasets

further shaping and joins on request …+8 more

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Where can I get historical box office data?

Datadory delivers it from four complementary records: Box Office Mojo's US chart archive with full gross histories for roughly 20,000 titles reaching the early 1980s, The Numbers' daily domestic ledger back toward its 1997 founding, the BFI's official weekly UK Top 15 series from 2017, and a 45,000-film budget-and-revenue bundle. Name the titles, markets and span; the sample arrives cut to exactly that scope.

How far back does box office history go?

Deeper than most projects scope, and unevenly so. US daily figures reach the early 1980s for major releases on the Box Office Mojo archive, The Numbers' daily charts run back toward its 1997 founding, the BFI's weekly UK series starts in 2017 with its Statistical Yearbook reaching 2002, and the Kaggle bundle spans 45,000 films released through July 2017.

Which dataset includes budgets alongside box office revenue?

Kaggle - The Movies Dataset (TMDB + MovieLens) puts budget and revenue on one row beside cast, crew, genres and keywords for 45,000 films, joined to 26 million MovieLens ratings. Kaggle - TMDB 5000 Movie Dataset is the lighter variant with 4,803 movies in two files. Both freeze at their upload vintages, so pair them with the live ledgers for anything newer.

Can I build commercial products on delivered box office rows?

Yes. Reuse rights travel with the rows: nothing in a Datadory delivery restricts republication or products built on them. The binding constraint is provenance rather than rights - figures are estimates stamped at specific periods, and cumulative totals move while a run continues - so every delivery carries period, vintage and estimate status as columns rather than footnotes.