Industry hub · Interactive media & services data provider

Interactive Media & Services Data: Video, Streaming, Music Catalogs and Social Graphs

Datadory covers interactive media & services data provider with 16 datasets spanning YouTube Data API v3 , Twitch Helix API and Wikimedia Dumps & Enterprise APIs . Delivered daily, weekly, or hourly — your call.

16 datasets 9 publishers Delivery: API, files, or your warehouse. Daily, weekly, or hourly.

The manifest

Every catch in this slice, ranked by our quality rubric — depth of documentation, freshness, breadth. Pick one, sample it, ship it.

#1 Google Developers

YouTube Data API v3

Coverage
Global - all public YouTube content, filterable per ISO…
#2 Twitch Developers

Twitch Helix API

Coverage
Global Twitch platform
#3 Datadory

Wikimedia Dumps & Enterprise APIs

Coverage
Global - all language editions and projects of…
#4 Spotify for Developers

Spotify Web API

Coverage
Global catalog with market availability expressed per…
#5

Stanford SNAP Large Network Collection

Coverage
Global online platforms - Facebook
#6 Tranco (KU Leuven DistriNet)

Tranco Top Sites Ranking

Coverage
Global domain popularity

Plus 10 more in this slice — each with its field dictionary and sample rows on its own page.

The lay of the water

Where interactive media & services data provider data actually comes from.

Pick your catch

Three to start with. The rest of the board is above.

Interactive media & services data provider Global - all public YouTube content… · Current-state snapshot of live…

YouTube Data API v3

kind · etag · id …+17 more

Interactive media & services data provider Global Twitch platform · Real-time current state of live…

Twitch Helix API

id · user_id · user_login …+16 more

Interactive media & services data provider Global - all language editions and… · Full edit history from each wiki's…

Wikimedia Dumps & Enterprise APIs

page.title · page.ns · revision.text …+7 more

Interactive media & services data provider Global catalog with market… · Current catalog state

Spotify Web API

id · name · duration_ms …+18 more

Interactive media & services data provider Global online platforms - Facebook · Snapshot windows roughly 2002 to 2020…

Stanford SNAP Large Network Collection

source_node_id · target_node_id · circle_name …+5 more

Interactive media & services data provider Global domain popularity · Daily lists since 2019

Tranco Top Sites Ranking

rank · domain

Interactive media & services data provider Global public web as crawled by the… · Captured fetches from 1996 to the…

Wayback Machine CDX Server API

urlkey · timestamp · original …+4 more

Interactive media & services data provider English-language Twitter · Static benchmark assembled from…

TweetEval Benchmark

text · label

Interactive media & services data provider Global Steam ecosystem · Rolling current state

SteamSpy — Ownership & Playtime Estimates

appid · name · owners …+10 more

Interactive media & services data provider Global - exchanges, facilities and… · Continuously updated by contributors

PeeringDB – Global Interconnection Database

record id · organization id / name · display name / aka / long name …+11 more

Interactive media & services data provider Global - volunteer clients in… · Continuous since February 2009

M-Lab Open Internet Performance Data

a.MeanThroughputMbps · a.MinRTT · a.LossRate …+7 more

Interactive media & services data provider Global - volunteer-run probes in… · Continuous collection since 2012

OONI Explorer – Global Internet Censorship Measurements

probe_cc · probe_asn · test_name …+9 more

Interactive media & services data provider Global - roughly 250 economies and… · Daily time series extending back to at…

APNIC Labs Measurements & Dashboards

date · cc · as …+5 more

Interactive media & services data provider Global - every routed ASN and… · Current-state snapshot of the global…

Hurricane Electric BGP Toolkit

ASN · Name · Prefix …+7 more

Interactive media & services data provider Global - every TLD delegated in the… · Current-state delegation snapshot

IANA Root Zone Database

domain · type · tld_manager …+1 more

Interactive media & services data provider Global, with regional commentary drawn… · Quarterly archive reaching back to Q1…

Domain Name Industry Brief & DNIB Data

quarter / reporting period · total_domain_name_registrations · com_net_domain_name_base …+3 more

Who fishes here

Frequently asked questions

Which datasets cover YouTube, Twitch and Spotify?
The three platform records anchor the slice. YouTube covers effectively all public videos with view counts, captions and region filters; Twitch serves realtime viewer counts, clips and creator analytics down to per-event pushes; Spotify exposes tracks, artists, albums, playlists and audiobooks including per-track audio features such as tempo and valence. All three arrive through Datadory with flattened, typed schemas.
Can I get historical time series for YouTube channels or Twitch viewership?
Yes, but understand the mechanics: platform surfaces report current state, so history is accumulated rather than queried retroactively. Datadory delivers those accumulations on your cadence - hourly around launch windows, daily for standard research loops - so the series builds in your warehouse instead of inside someone else's quota window. Archived Wayback captures add before-and-after evidence for specific pages.
Are there social network graph datasets large enough for benchmarks?
The Stanford SNAP Large Network Collection packages 80-plus graphs drawn from Facebook, Twitter/X, Google+, YouTube, Reddit, Twitch and LiveJournal, ranging from thousands of edges to com-Friendster's 1.8 billion. Files ship as edge lists and adjacency lists that load into any graph engine unmodified, which is why they remain the standard benchmark substrate for community detection.
Is there labeled Twitter data for text classification?
Yes. TweetEval Benchmark fixes 200,785 tweets - roughly 14.2 MB - across seven tasks: irony, hate speech, offensive language, stance, emoji, emotion and sentiment, each with frozen train/validation/test splits. Fixed splits are the point: results become comparable to published baselines instead of re-partitioned approximations.
How far back does website history data go?
To 1996 for the web broadly: the Wayback Machine CDX index covers hundreds of billions of captures, and the earliest record observed in this slice is netflix.com in January 1999. Each capture record returns urlkey, timestamp, mimetype, statuscode, digest and length, filterable by URL pattern, date range and mimetype - enough to diff a competitor's page across a decade.
Does the slice include internet performance and censorship measurements?
Yes. M-Lab contributes petabytes of speed-test and traceroute records continuous since February 2009, OONI Explorer adds hundreds of millions of censorship measurements from probes in nearly 200 countries since 2012, and APNIC Labs publishes daily IPv6, DNSSEC, RPKI and QUIC adoption series per country and ASN back to at least October 2013. Together they connect audience behavior to how the network actually performs.
What does a Datadory sample include?
Real rows, not screenshots: a video-channel extract limited to the creators and fields you specify, a week of viewer-count series at your chosen interval with the field dictionary attached, an edge-list slice sized to load in your environment, or a capture-history pull for named domains. Coverage statements and sample data are yours to keep whether or not you continue.

See the rows before you commit

Any catch on this board, sampled against your own question. API, files, or your warehouse. Daily, weekly, or hourly..