Broadcasting · TheTVDB (Whip Media)
TheTVDB - TV and Film Metadata Database
Datadory delivers TheTVDB tv and film metadata database data covering every series, season, episode, movie, person, company and award record its contributor community curates - titles, aliases, air dates, runtimes, statuses, popularity scores, artwork and translations in dozens of languages - the catalogue media-centre platforms like Plex and Kodi are built on.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Global - country and language fields sit on every series, and name/overview translations span dozens of codes from ara to zho
- How far back
- Full historical archive of aired television and film, forward-looking nextAired values where schedules are known, and a change feed recording edits and merges
- How fine
- Series, season, episode, artwork, person, company, award and list level - one record per entity, joined by stable identifiers
What is the TheTVDB TV and film metadata database?
TheTVDB - TV and Film Metadata Database is the Broadcasting catalog's descriptive-layer record: a fan-contributed, community-curated catalogue operated by Whip Media that describes itself as the most accurate source for TV and film, and that pushes everything contributors add out to other sites, mobile apps and devices. Its surface covers series, seasons, episodes, artwork, movies, people, companies, awards, genres, languages, countries, content ratings and lists, plus a tagging layer that classifies how records relate. This is the metadata underneath media centres - Plex and Kodi are built on it.
What makes the catalogue distinctive is curation depth at the edges of television: aliases and translated names for regional broadcasts, absolute-number ordering for shows watched in production order, finale markers so end-of-run features need no heuristics, and a change feed that reports merges when duplicate records collapse.
Get a sample of this dataset to see series and episode records laid against your use case.
What do sample rows look like?
Two series records and one episode, shaped exactly as they deliver:
# SeriesBaseRecord
id : 121361
name : The Thick of It
slug : the-thick-of-it
country : GB originalLanguage: eng
status : Ended
firstAired : 2005-05-19 lastAired: 2012-10-20
averageRuntime : 29 score : 4.2131
# SeriesBaseRecord
id : 73141
name : Mystery Science Theater 3000
country : US status: Ended
firstAired : 1988-11-24 lastAired: 1999-08-08
averageRuntime : 91
# EpisodeBaseRecord (seriesId 121361)
aired : 2009-11-21 runtime: 29
seasonNumber : 3 number : 1
absoluteNumber : 17 finaleType: (null)Read the anatomy rather than the individual titles. A British political satire with a 29-minute average runtime and a clean Ended status; an American show whose 91-minute average runtime immediately separates it from standard half-hour comedies; and an episode placed three ways at once - season 3, episode 1, absolute 17 - which is precisely the redundancy that lets ordering logic survive re-edits and special episodes. seriesId joins the episode back to its parent without a lookup table.
What fields does the dataset include?
Sixteen documented fields carry every delivered row - eight describing a series, five placing an episode inside it, three carrying curation state. A further group (overview translations, artwork references, ordering flags, merge targets) is defined across the same records but was not value-sampled during the cataloging pass, so those fold under additional fields on request.
How far does coverage reach?
- Geography: Global.
countryandoriginalLanguagesit on every series record, and name and overview translations span dozens of language codes, so a Korean drama's French release resolves to the same series as its Seoul broadcast. - Temporal: The full historical archive of curated television and film, plus forward
nextAiredvalues where schedules are known, and a change feed recording edits and merges over time. - Granularity: Series, season, episode, artwork, person, company, award and list level - one stable identifier per entity, so artwork sets and cast lists hang off the same key as the title.
That combination - historical depth, forward dates and entity-level grain - is what makes this the reference catalogue for episode guides rather than a schedule grid or a ratings panel. For actual audience measurement the neighbours are BARB Weekly Top 50 Shows and Viewing Data in the UK and the Internet Archive TV News Archive for caption-searchable news history.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Every delivery ships the complete field dictionary above, the sample rows and the coverage profile mapped to your scope - series only, one franchise, or the whole catalogue. Because records carry a last-edited timestamp and merges surface through the change feed, teams typically take the historical backfill once and keep new and edited records rotating in thereafter, with no silent drift between syncs.
Who uses this data, and for what?
- Media-centre and product developers build guides and discovery features on stable identifiers: series ID joins artwork, cast and episode lists; translations localize the interface without re-sourcing titles.
- Content-acquisition and programming analysts profile franchises - season counts, runtimes, finale structure, country and language mix - before licensing conversations, using curation rather than box-office proxies.
- Market researchers studying global content flows read country-of-origin and translation coverage as a footprint of international distribution, show by show.
- Journalists and academics cite structured facts about programmes - first and last air dates, episode counts, ordering - instead of scraping fan wikis of unknown provenance.
Across all four, the constant is that description travels with identification: one id resolves title, aliases, translations, artwork and episode tree together.
Which personas get the most value?
Developers and data-product builders get the canonical join layer for anything that displays television - the reason media-centre platforms standardized on it. Data scientists and ML engineers get labeled, translated text and artwork references for recommendation and content-matching models. Market researchers and consultants get programme-level descriptors that make international catalogues comparable. Competitive intelligence and product teams track rivals' catalogue shapes by genre, country and runtime mix. Journalists, academics and students cite community-audited facts with stable identifiers behind them. Persona workflows live at data scientists x broadcasting, developers & data-product builders x broadcasting and journalists, academics & students x broadcasting.
How does it compare to alternatives in its slice?
Within TV metadata, the trade-offs run on curation model and ordering fidelity. TVmaze is the realtime schedule-first option: unauthenticated JSON, full future listings, and a share-alike licence - the head-to-head lives on our comparison page. TMDB's TV section brings daily bulk exports and deep artwork but restricts free keys to non-commercial use. TheTVDB's edge is ordering and localization discipline - absolute numbers, airs-before flags and the deepest alias/translation habit in the slice - which is why media centres lean on it.
The usual pattern is to stack them: TVmaze for schedules, TMDB for bulk artwork pulls, TheTVDB as the identity spine that keeps both consistent. Where these rank for the industry sits on our broadcasting datasets ranking.
What should I know before requesting a sample?
Four notes worth having in hand:
- Community-curated, professionally consumed - Contributors add and correct the records; media-centre platforms distribute them. Curation depth is the product, which is why aliases and translations arrive populated rather than null.
- Three ways to number an episode - Season number, absolute number and airs-before/after flags coexist on one episode record, so anime seasons, specials and split cour order correctly without a second table.
- Merges leave a forwarding address - When duplicate records collapse, the change feed carries a pointer to the surviving record - incremental syncs keep their joins intact instead of silently dropping history.
- Cross-identifier lookup - Records resolve against external identifiers used by other platforms, so TheTVDB IDs map onto IDs your systems already hold.
The catalogue is community-curated, so coverage concentrates where contributors concentrate: mainstream English-language television is exhaustive, while some regional or very recent programming carries thinner translation sets until someone adds them. Exact catalogue size is not published, which is why the sample - cut to the series, genres or countries you name - settles scope questions faster than any spec sheet. Where this set stops, neighbours pick up: TMDB for daily bulk exports, TVmaze for realtime schedules, and BARB for who actually watched.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
id | integer | TheTVDB identifier for the series or episode - the stable key every alias, artwork and update record joins through. | 121361 |
name | string | Primary title of the series or episode as the community curates it. | The Thick of It |
slug | string | URL-safe identifier for the record, derived from the title. | the-thick-of-it |
aliases | enum | Alternative titles in other languages or regions, so a Japanese broadcast of a British comedy resolves to the same series. | ティック・オブ・ミット (ja) |
firstAired | date | Date the series first aired anywhere. | 2005-05-19 |
lastAired | date | Date the most recent episode aired. | 2012-10-20 |
nextAired | date | Date of the next scheduled episode where one is known - the field EPG builders watch. | <returned in your sample> |
averageRuntime | integer | Average episode runtime in minutes across the series. | 29 |
status | enum | Series production state such as Continuing or Ended. | Ended |
score | number | Relative popularity score computed over the catalogue. | 4.2131 |
country | string | Country of origin for the series. | GB |
originalLanguage | string | Original language of production. | eng |
nameTranslations | enum | Language codes for which translated names exist - the multilingual spine of the catalogue. | ara, chi, dan, deu, fra, jpn |
seasonNumber | integer | Season number the episode belongs to. | 3 |
absoluteNumber | integer | Episode number counted continuously across all seasons, for shows watched in production rather than broadcast order. | 24 |
finaleType | enum | Marks season or series finale episodes, so an end-of-run feature needs no heuristics. | series |
Questions buyers ask
What does each TheTVDB series record contain?
An identifier, primary title and slug, alternative titles and their language codes, country and original language of production, first, last and next air dates, average runtime, production status, a relative popularity score, and links into seasons, episodes and artwork. Episodes add season, absolute and airs-before numbering plus finale markers.
How many records does the database cover?
The catalogue spans hundreds of thousands of series and movies across entities - series, seasons, episodes, artwork, people, companies, awards and lists - curated continuously since the mid-2000s. Exact totals are not published, so samples are cut to named series or genres to settle scope.
Are translations included for non-English programming?
Yes. Every record carries language codes for which translated names and overviews exist, spanning dozens of languages. A single series identifier therefore resolves its German, Japanese and Spanish presentations without joining a separate localisation table.
How are episodes ordered for shows with specials or multiple seasons?
Three schemes coexist on the episode record: the season number, the continuous absolute number, and airs-before or airs-after placement flags. Anime seasons, specials and split cours therefore sort correctly under whichever scheme your product uses.
Which products are built on this metadata?
Media-centre platforms including Plex and Kodi draw on TheTVDB for series, episode and artwork information, which is why its identifiers behave as a de facto join layer across home-theatre and discovery tooling.
Can a sample be cut to specific series or genres?
Yes. Name the series, franchises, genres or countries and the sample arrives shaped to that scope with the complete field dictionary attached. Samples precede any commitment, and the schema in the sample is the schema you ship against.
What is this dataset best used for?
Anything that has to describe television consistently: episode guides, discovery rails, catalogue enrichment, recommendation features and cross-platform identity mapping. It describes programmes; it does not measure audiences - pair it with a viewing panel for that.
Notes on this record
- Community-curated, professionally consumed Contributors add and correct the records; media-centre platforms distribute them. Curation depth is the product, which is why aliases and translations arrive populated rather than null.
- Three ways to number an episode Season number, absolute number and airs-before/after flags coexist on one episode record, so anime seasons, specials and split cour order correctly without a second table.
- Merges leave a forwarding address When duplicate records collapse, the change feed carries a pointer to the surviving record - incremental syncs keep their joins intact instead of silently dropping history.
- Cross-identifier lookup Records resolve against external identifiers used by other platforms, so TheTVDB IDs map onto IDs your systems already hold.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.