Internet Archive TV News Archive Data

Datadory delivers internet archive tv news archive data covering approximately 4.37 million US and international television news broadcasts dated from March 2001 through the present - each record carrying an identifier encoding channel and airtime, a programme-channel-date title, broadcast date, ingest timestamp and description, searchable at closed-caption level down to half-hour segments, delivered daily, weekly, or hourly.

What is the Internet Archive TV News Archive?

It is the Internet Archive's broadcast-television collection: roughly 4.37 million US and international news broadcasts as of August 2026, each indexed with its closed-caption transcript so a phrase said on air retrieves the exact programme that said it. Coverage opens in March 2001 and runs continuously to the present - launch depth came from US national networks plus San Francisco and Washington DC locals, a donation of about 40,000 tapes from Marion Stokes's estate pulled in Philadelphia and Boston, and international broadcasters including M1 (Hungary), Press TV, Globovision, Syrian News and Zvezda now sit in the same index.

Within Datadory's fifteen-dataset broadcasting slice it holds a unique position: schedule catalogues describe programming, ratings currencies measure audiences, and this is the only product that preserves what news broadcasts actually said, searchable at caption level across a quarter-century. It scores 9 out of 10 on Datadory's quality rubric, with field definitions verified against sampled records spanning the collection's full range.

Get a sample of this dataset and see your channels, dates and phrases resolved against live rows.

What do sample rows from the dataset look like?

One row per broadcast programme, exactly as it arrives:

identifier : CNN_20010314_050000_Larry_King_Live
title      : Larry King Live : CNN : March 14, 2001 5:00am-6:00am EST
date       : 2001-03-14

identifier : KQED_20260821_000000_BBC_News_The_Context_USA
title      : BBC News The Context USA : KQED : August 21, 2026 12:00am-12:30am PDT
date       : 2026-08-21

identifier : M1_20251130_180000_Idojaras-jelentes
title      : Időjárás-jelentés : M1 : November 30, 2025 7:00pm-7:05pm CET
date       : 2025-11-30
publicdate : 2025-11-30T18:00:00Z
description: Részletes időjárás-jelentés.

These are three real records spanning the shelf, and they make the structural argument better than any description could. A CNN hour from March 2001, a KQED airing of BBC News The Context USA from August 21, 2026 and a five-minute Hungarian weather report on M1 all land in one identical shape - which means a fact-checker verifying a 2001 claim, a media analyst tracking this month's cycle and an NLP team mining non-English captions query the same dictionary. Note the identifier doing quiet work: M1_20251130_180000_Idojaras-jelentes encodes channel, date and time before you read another field, so keys stay human-auditable even at four million rows.

What fields does the dataset include?

Six verified fields carry every one of the roughly 4.37 million records, defined in the dictionary below. None of them is decorative. The identifier is a structured key rather than a random string, so channel and airtime survive even when titles change format; the pair of dates - broadcast versus ingested - is what lets you separate what aired from when it became retrievable.

Additional fields on request. The wider delivery carries more than the core row shape: caption text aligned to the video timeline for teams who need the transcript column beside the metadata, streaming clip references for embed workflows, borrow access flags on full broadcasts, and derived cuts such as per-channel or per-day counts. Each gets defined and exemplified against live records before your pipeline locks a schema, so the dictionary you buy is the dictionary you tested.

Where does coverage reach across geography, time and granularity?

  • Geo: Primarily United States national and local news, plus international broadcasters. Earliest holdings centre on San Francisco and Washington DC stations; the Marion Stokes donation added Philadelphia and Boston; recent results surface M1 (Hungary), Press TV, Globovision, Syrian News and Zvezda.
  • Temporal: Records dated from March 2001 through the present, ingested continuously - a quarter-century of news cycles in one consistent frame, new footage landing days old beside material twenty-five years old.
  • Granularity: One record per broadcast programme, typically half-hour to hour-long segments, searchable at caption-text level rather than only by title or channel.

That footprint is the pitch: schedule catalogues describe what was meant to air, ratings currencies count who watched, and this archive preserves what was actually said - searchable by the words in the captions.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the slice and set the cadence to match the decision being fed. Fact-checking desks want the newest broadcasts rotating in fast, because a claim dies inside a news cycle; historical researchers take the March 2001-onward backfill once and append as it grows; ML teams pin a fixed snapshot for reproducible training runs and refresh on their own clock. Because broadcast date and publicdate ride on every record, incremental pulls key cleanly without double-counting. The sample comes first, so channels, span and cadence are settled facts before any commitment.

Who builds on it, and for what?

  • Journalists and academics run claim-checking as a lookup: quote in, moment-on-air out, with citable footage behind every hit. The journalists academics use cases page shows the workflow against the rest of the broadcasting shelf.
  • Competitive intelligence and product teams watch when a rival, a category or a phrase entered the news cycle and how coverage curved afterward. The competitive intel product teams use cases page covers the patterns.
  • Data scientists mine a quarter-century of aligned video-transcript pairs for media-measurement and NLP work - the data scientists use cases page has the model-side detail, and ml model training shows the broader pattern.
  • Market researchers quantify coverage themes for client briefs with a primary-source corpus instead of a clipping service; citation-grade research is the general case.

Persona fit is lopsided by design: sales teams find no prospecting signal in broadcast captions, and e-commerce operators find no pricing or assortment content - the archive is unambiguous about what job it does.

Which notes and neighboring datasets pair with it?

Three notes worth having before you request a sample. First, this is the news slice of the Internet Archive, not the whole non-profit: the Wayback Machine's web captures and the Digital Library's scanned texts are separate collections with separate dictionaries, catalogued under their own industries. Second, full broadcasts travel under a borrowing model while short clips stream freely - if your workflow needs embeddable moments, scope the clip layer at sample time; if it needs whole programmes, say so, because the two arrive differently. Third, the station roster is not published as a list - it is derivable from the records themselves, which is precisely what a scoped sample settles: name the channels you assume you need and the sample confirms them against live rows.

Notes that pair well with this page:

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary — six verified fields, one row per broadcast programme
fieldtypedefinitionexample
identifierstringUnique Internet Archive item identifier, structurally encoding channel, date and time of broadcast.CNN_20010314_050000_Larry_King_Live
titlestringHuman-readable title combining programme name, channel and air datetime.Larry King Live : CNN : March 14, 2001 5:00am-6:00am EST
channelstringBroadcast channel or station call sign the programme aired on.CNN
datedateBroadcast date of the programme.2001-03-14
publicdatedatetimeDate and time the item entered the Internet Archive - the ingest timestamp distinct from the air date.2025-11-30T18:00:00Z
descriptiontextItem description or programme summary where supplied.Részletes időjárás-jelentés.

Coverage chips — geography, time, granularity

dimensioncoverage
GeographicPrimarily United States national and local news plus international broadcasters; earliest holdings include San Francisco, Washington DC, Philadelphia and Boston stations
TemporalMarch 2001 through the present, ingested continuously
GranularityIndividual broadcast programmes (typically half-hour to hour-long segments), searchable at caption-text level

Questions buyers ask

How far back does the internet archive tv news archive data reach?

Holdings are dated from March 2001 through the present. The earliest sampled record is a Larry King Live broadcast on CNN from March 14, 2001, carrying identifier CNN_20010314_050000_Larry_King_Live. Launch coverage in September 2012 targeted US national networks plus San Francisco and Washington DC local stations before donations widened the base.

What fields does each broadcast record include?

Six verified fields per item: an identifier encoding channel, date and time; a title combining programme name, channel and air datetime; the channel or station call sign; the broadcast date; publicdate, marking when the item entered the Archive; and a description where supplied. Every field is pinned against live records in the sample before you commit.

Does coverage extend beyond United States networks?

Yes, though the core remains US national and local news. Sampled records include M1 Hungary airing a Hungarian-language weather report, and recent results surface Press TV, Globovision, Syrian News, Zvezda and KQED alongside CNN. Caption text makes international material retrievable by transcript, not just by channel or date.

What can I actually do with four million caption-searchable broadcasts?

Fact-check a claim to the moment it was said on air; measure when a competitor, product or phrase entered the news cycle and how coverage evolved; train language models on a quarter-century of aligned video-transcript pairs; or quantify coverage themes for briefs with citable primary-source footage behind every row.

How granular is one record?

One record is one broadcast programme segment, typically half-hour to hour-long, searchable at caption-text level. A five-minute M1 weather report sits beside a full hour of Larry King Live in the same shape, so short updates and long-form programmes filter on identical fields.

Can a sample be scoped to my channels, dates or topics?

Yes - that is what the sample is for. Name the channels, the date span and the phrases you care about, and real rows come back cut to that scope with the six-field dictionary unchanged, so you validate the exact extract your pipeline will receive rather than a generic preview.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing