Movies & Entertainment · IMDb
IMDb Non-Commercial Datasets
Datadory delivers imdb non commercial datasets data covering every movie, TV series, short, video and video game IMDb indexes - twelve million-plus titles and sixteen million-plus people joined on stable tconst and nconst identifiers, with crowd ratings, localized alternate titles, full crew, top-billed cast and episode-to-series linkage carried alongside.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Global: every film, television series, short, video and video game production indexed by IMDb worldwide.
- How far back
- From cinema's beginnings to currently releasing titles. Snapshots land complete; you choose the cadence they arrive on.
- How fine
- Title-level and person-level records, with credit-level principals underneath and episode-level linkage tying installments to their parent series.
What is the IMDb Non-Commercial Datasets dataset?
It is the reference catalog of screen entertainment, delivered as seven coordinated tables instead of one sprawling dump. Title basics carries every movie, short, TV series, episode, video and video game IMDb indexes - twelve million-plus titles - with its tconst, type, year, runtime and up to three genres. Around it sit ratings, alternate titles, crew, principals and episodes, plus a people table covering sixteen million-plus individuals.
The design decision worth admiring: everything joins on two stable identifiers, tconst for titles and nconst for people. A rating finds its title, a credit finds its person, an episode finds its series, and none of it requires fuzzy name matching. Datadory normalizes the whole set into query-ready form, so "what is this title and who made it" becomes a column lookup - see how movies & entertainment data families stack.
What do sample rows look like?
One row per title in basics, with the rating arriving on the shared identifier:
# one row per title in title.basics; ratings, localized
# titles, crew, principals and episodes join on tconst
file : title.basics
tconst : tt0111161
titleType : movie
primaryTitle : The Shawshank Redemption
startYear : 1994
runtimeMinutes : 142
genres : Drama
# the same tconst carries its crowd score:
file : title.ratings
averageRating : 9.3
numVotes : 2900000
# people arrive on their own nconst; episodes tie back
# to a parent series with season and episode numbers;
# a backslash-N marks a missing value everywhereThe pattern worth noticing: each table holds one grain and refuses to mix them, so the catalog stays flat while the depth lives in the joins. Rows reflect the August 2026 review pass; a fresh pull arrives current to the day it ships.
Which fields does each record include?
Thirty fields span the seven tables, verified against the published dictionary during the research pass. The title spine: tconst, titleType, primaryTitle, originalTitle, isAdult, startYear, endYear, runtimeMinutes and genres. Scoring adds averageRating - a weighted average of individual user ratings - and numVotes. Alternate titles contribute region, language, types, ordering and isOriginalTitle, so one work's life under other flags becomes filterable.
Credits and people complete it: directors and writers as nconst arrays, then nconst, category, job, characters and ordering per principal credit. Episodes tie installments to their parents with parentTconst, seasonNumber and episodeNumber, and name basics rounds out with primaryName, birthYear, deathYear, primaryProfession and knownForTitles. The full dictionary follows in the table below.
What does coverage look like across geography, time and granularity?
Geography - global by construction: every film, television series, short, video and video game production IMDb indexes worldwide, whatever country it came out of.
Temporal - from cinema's beginnings to currently releasing titles, with end years marking closed runs. Snapshots land complete whenever they ship, and the cadence they reach you on is yours to set - which turns a moving catalog into a panel the moment you accumulate pulls. Pair the slice with GroupLens MovieLens Datasets for interaction history or Box Office Mojo for grosses when a question needs behavior or money next to the record.
Granularity - title-level and person-level records, with credit-level principals underneath and episode-level linkage tying installments to their series. Three grains, three tables, no ambiguity.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Who uses this data, and for what?
- Recommendation and personalization systems - weighted ratings and vote counts on stable identifiers remain the canonical cold-start signal for recommender models; see ml model training.
- Catalog enrichment - year, runtime, genres, cast and localized titles attach to any corpus you already own via a one-hop join.
- Localization research - alternate titles by region, language and type show how a single work travels under different names market to market.
- Talent and credit analysis - principals with categories and roles, plus professions and known-for titles per person, reconstruct careers credit by credit.
- Demand proxies - vote volume complements box office and streaming signals for a fuller interest picture; see demand forecasting.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
tconst | string | Alphanumeric unique identifier of the title. | tt0111161 |
titleType | enum | Type or format of the title: movie, short, tvSeries, tvEpisode, video, videoGame and more. | movie |
primaryTitle | string | The more popular title used for promotional purposes. | The Shawshank Redemption |
originalTitle | string | Original-language title of the work. | Original-language release title |
isAdult | boolean | 0 or 1 flag marking adult titles. | 0 |
startYear | integer | Release year; for TV series the series start year. | 1994 |
endYear | integer | TV series end year; the missing-value marker for all other types. | \N |
runtimeMinutes | integer | Primary running time in minutes. | 142 |
genres | text | Up to three genres associated with the title, comma-separated. | Drama |
averageRating | number | Weighted average of all individual user ratings on IMDb. | 9.3 |
numVotes | integer | Number of votes the title has received. | 2900000 |
ordering | integer | Ordering rank within the source list, for aka locales and principal credits alike. | 1 |
region | string | ISO 3166 region code for the localized title variant. | US |
language | string | Language code of the localized title. | language code per locale |
types | text | Alternate-title type set: alternative, dvd, festival, tv, video, working, original, imdbDisplay; new values may appear. | imdbDisplay |
isOriginalTitle | boolean | 0 or 1 flag whether the locale's title is the original title. | 1 |
directors | string | Comma-separated array of nconst identifiers for directors. | director nconsts, comma-separated |
writers | string | Comma-separated array of nconst identifiers for writers. | writer nconsts, comma-separated |
nconst | string | Alphanumeric unique identifier of the name or person. | nm0000001 |
category | string | Job category the person was credited in for the title. | director |
job | string | Specific job description if applicable, otherwise the missing-value marker. | \N |
characters | string | Character name or names played if applicable, otherwise the missing-value marker. | played-character names |
parentTconst | string | tconst of the parent TV series for an episode. | parent series tconst |
seasonNumber | integer | Season number the episode belongs to. | 1 |
episodeNumber | integer | Episode number within the season. | 1 |
primaryName | string | Most commonly credited name of the person. | most-credited name |
birthYear | integer | Birth year of the person; the missing-value marker if unknown. | four-digit year |
deathYear | integer | Death year of the person; the missing-value marker if living or unknown. | four-digit year |
primaryProfession | text | Top three professions of the person, comma-separated. | three professions, comma-separated |
knownForTitles | text | Array of tconsts the person is known for. | known-for tconsts |
What teams do with it
- Recommendation and personalization systems Ratings and vote counts on stable title identifiers are the canonical cold-start signal for recommender models.
- Catalog enrichment Attach year, runtime, genres, cast and localized titles to any title corpus by joining on tconst.
- Localization research Alternate titles by region, language and type show how one work travels under different names.
- Talent and credit analysis Principals with categories, jobs and characters, plus professions and known-for titles per person, map careers credit by credit.
- Franchise and series tracking Episode-level linkage to parent series with season and episode numbers makes serial structures queryable.
- Demand proxies Vote volume and weighted averages complement box office and streaming signals for a fuller interest picture.
Questions buyers ask
How many titles and people does the dataset cover?
Twelve million-plus titles and sixteen million-plus people uncompressed across the seven tables, roughly two gigabytes compressed in aggregate, with the principals table the largest single piece. Coverage runs from cinema's beginnings through currently releasing titles, worldwide.
What is one row in this dataset?
It depends on the table: one row per title in basics, one per title-rating pair, one per localized alternate title, one per principal credit, one per episode, and one per person in name basics. Aggregate at the grain your question needs before you count.
Which identifier links the tables together?
Titles travel on tconst and people on nconst. Ratings, alternate titles, crew, principals and episodes all reference tconst, while crew, principals and name basics reference nconst, so any two tables join on at most one hop. See our guide to IMDb identifier cross-linking.
Do ratings come with the titles?
Yes. Each rated title carries a weighted average of individual user ratings plus the vote count behind it, delivered on the same identifier as the title row, so scoring joins to catalog attributes without reconciliation work.
How are TV series handled compared to movies?
Series appear as their own title rows, and each episode arrives as a row linking back to the parent series with season and episode numbers. Series rows also carry an end year, which films leave marked missing.
Can a sample be scoped to my slice?
That is what the sample is for. Name the title types, years, genres or regions you care about and real rows come back shaped like your production tables, including the columns folded out beyond the core dictionary.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.