Datadory notebook
The Interactive Home Entertainment Data Guide (2026)
Datadory delivers interactive home entertainment data covering every layer the industry gets asked about: 261,285 storefront applications with nested price packages, month-by-month player histories reaching each title's own launch, estimated owner ranges sitting beside raw review counts, 500,000-plus games carrying 1.1M ratings and machine-ranked look-alikes, weekly console retail charts across seven regions back to December 2004, 1,509,534 indie listings, and hundreds of thousands of adjudicated speedrun runs - typed rows, delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
The Interactive Home Entertainment data landscape in 2026
This is the one industry in the catalog where the customers generate the dataset as a byproduct of consuming the product. Every session leaves a player count, every purchase leaves a price point, every leaderboard attempt leaves an adjudicated time. The result is measurement depth consumer-packaged-goods analysts would trade a limb for - and it shows in the pool: Datadory catalogs 18 datasets here, 16 primary records plus 2 related neighbors, scoring 10 down to 4 on the quality rubric with a mean of 7.75 against a 7.81 catalog-wide average.
Four Steam records anchor the middle of the stack. The Steam Web API serves current-player snapshots, global achievement percentages and per-user owned-games playtime - verified live in August 2026 at 871,553 concurrent players on a single flagship app. The Steam Store catalog carries 261,285 searchable applications with prices, discounts, genres and review summaries. SteamSpy estimates ownership and playtime per application, with aggregates accumulating since March 2009, and Steam Charts keeps month-by-month average-player tables reaching back to each title's own launch.
Around them sit the specialists. RAWG documents more than 500,000 games across 50 platforms in 31 typed fields; itch.io tracks 1,509,534 indie listings with an archive running to its 2013 launch; VGChartz indexes weekly console retail charts across seven regions back to December 2004 beside a 67,171-title database; the Speedrun.com REST API logs hundreds of thousands of verified runs; and TwitchTracker ranks games and channels by viewership with monthly histories stretching back years.
Five pooled neighbors frame the market rather than measure it: Google Trends supplies normalized search interest down to metro level, GDELT reads the news cycle in 100-plus languages back to January 1979, Common Crawl captures storefronts and wikis at web scale across 300 billion-plus pages, the Chrome UX Report benchmarks portal performance back to 2017, and the Facebook Graph API carries publisher social presence. Two related streaming records - the Twitch Helix API and YouTube Data API v3 - round the pooled total out to 18.
Every record behind this guide documents its own field dictionary, sample rows and coverage span, so the grain is visible before anything lands in a model. Browse the full pool on the interactive-home-entertainment data hub.
The core datasets to know
Sixteen primary records sort into five working groups, and knowing which group owns a question decides which one you pull.
Live engagement
- Steam Web API (quality score 9). Valve's authoritative telemetry: per-application app lists, news feeds, global achievement unlock percentages and current-player snapshots, plus owned-games and playtime records for consenting accounts. The August 2026 research pass captured 871,553 concurrent players on appid 570 and 40.2 percent of all players having earned Counter-Strike 2's headline achievement - a demand gauge and a completion curve in two cells. Player figures arrive point-in-time; history is built by capturing on a schedule.
- Steam Charts - Concurrent Players (quality score 7). The history layer the snapshots lack: one row per game per month back to each title's own launch, alongside live standings, 24-hour and all-time peaks. A 2011 release carries over a decade of engagement curve; Counter-Strike 2's February 2025 row reads 1,039,662 average players, up 3.60 percent on the month, peaking at 1,818,368.
- TwitchTracker - Game & Channel Rankings (quality score 7). The streaming side: average and peak concurrent viewers, hours watched, minutes broadcast, follower gains and share-of-platform percentages per game and per channel, on trailing seven-day leaderboards, thirty-day summaries and multi-year monthly histories containing the platform's all-time peak month of May 2021.
Ownership and money
- SteamSpy API (quality score 8). The only record here that estimates who owns a game: owner ranges such as 100,000,000 .. 200,000,000 sit beside observed positive and negative review tallies - 7,642,084 against 1,173,003 for Counter-Strike: Global Offensive, an 87-percent positive ratio computed from counts rather than badges. Average and median playtime ship on both lifetime and trailing-two-week bases, accumulating since March 2009, and prices are quoted in US cents.
- VGChartz - Video Game Sales Charts (quality score 6). Retail units, weekly, across Global, USA, Europe, UK, Germany, France and Japan, indexed back to December 2004, with lifetime platform hardware totals - PlayStation 2 at 160.01M global units, Nintendo Switch at 155.32M - tie ratios and a 67,171-title database. Twenty documented fields per row, and every figure labeled an estimate by the publisher itself.
- Video Game Sales with Ratings (Kaggle) (quality score 7). The frozen benchmark: 16,598 titles selling above 100,000 copies, released 1980 through 2016, one row per title per platform with unit splits for North America, Europe, Japan and other regions. Japan is broken out separately because it behaves differently - handheld-first franchises and home-console RPG lineups produce regional skews a two-region model would erase.
Catalog and metadata
Community and competition
- Speedrun.com REST API (quality score 8). Competitive completion as an adjudicated record: hundreds of thousands of verified runs across tens of thousands of games, 42 documented fields per run, three timing methods carried side by side down to the millisecond, video evidence and runner context. Performed-date and verify-date ride on every run, which turns moderation latency into a measurable quantity, and a
system.emulatedboolean keeps emulator-based analyses honest.
Attention, news and web context
- Google Trends (quality score 7): the normalized 0-100 search-interest series for any term, worldwide down to US metro level, back to 2004, with top-25 and rising-query ranks across 210 US DMAs and roughly fifty countries. Values are relative to the comparison set, never absolute volume.
- GDELT Project 2.0 (quality score 9): CAMEO-coded news events between named actors with tone and Goldstein scores, georeferenced to city centroids in 100-plus languages, from January 1, 1979 onward, plus a knowledge graph codifying people, organizations and themes per document - a quarter-billion-plus event records and roughly 2.5 TB of graph per year.
- Common Crawl Web Corpus (quality score 10, the pool's highest): 300 billion-plus pages across 94-plus named releases since 2008, one record per fetched URL with capture timestamp, content type and digest. Each release is preserved rather than overwritten, so what a store page said in any given month stays answerable deterministically.
- Chrome UX Report (CrUX) (quality score 9): traffic-weighted Core Web Vitals histograms and p75 LCP, INP and CLS per origin, split across desktop, phone and tablet, with a rolling 28-day view and monthly history back to 2017.
- Facebook Graph API (quality score 8): Page, Post, Comment and Photo nodes with each Page's fan_count, followers_count and talking_about_count trio and posts carrying message text with ISO 8601 timestamps. Live nodes, no historical backfill.
- Reddit Data API via Pushshift (quality score 4) rounds out the group with comment and submission records, scoped to community-moderation use - the pool's lowest score, and the honest floor of what this slice reaches.
Read together, the sixteen cover a complete research arc - demand, ownership, catalog, competition and context - and the two related streaming records extend the engagement layer onto live video.
Interactive Home Entertainment data by use case
Ten recurring jobs define how teams put this pool to work, and each maps to named records:
- Estimate unit sales and install bases. Triangulate SteamSpy owner-range midpoints against VGChartz weekly software figures and lifetime platform totals, pinned to the 16,598-title Kaggle CSV as a fixed historical benchmark. Carry both bounds of every owners band instead of collapsing to a number.
- Monitor whether a title is growing or fading. Layer Steam Web API current-player snapshots over Steam Charts monthly averages and TwitchTracker hours-watched rankings - three independent views of the same question across PC and streaming audiences.
- Build recommendation and discovery models. Train on RAWG's 1.1M ratings and 58K tags with its machine-generated similar-game suggestions as ground truth, then enrich with SteamSpy tag vectors and playtime aggregates as behavioral features.
- Track PC price movements and discount depth. Steam Store records nest package tiers and discount state inside each application, with prices localized across supported currencies; SteamSpy's US-cent price fields cross-check the same titles from a second angle.
- Profile the indie long tail. Cut itch.io's 1,509,534 listings by price band, tag and platform target - coverage storefront-level trackers never reach, with jam weekends showing up as visible spikes that make event-driven analysis possible without stitching sources.
- Run competitive speedrunning analytics. Group verified Speedrun.com runs by game, category, level and leaderboard; compare timing methods within a ruleset and compute reporting lag from performed-versus-verified timestamps.
- Gauge pre-launch demand. Google Trends' 0-100 interest series by country, state and metro doubles as an attention proxy for upcoming releases and genre shifts; rising-query ranks catch breakouts before volume does.
- Benchmark web performance for game portals. Compare CrUX p75 LCP and INP histograms for storefront and launcher origins across monthly tables going back to 2017, split by device class.
- Build alternative quant signals. Join GDELT event and knowledge-graph rows with Steam player-count series to model news sentiment around individual publishers, and use Common Crawl to read the wikis and forums where the verdict actually gets written.
- Support journalism and academic study. Chart archives, adjudicated world records and the fixed sales benchmark are citable institutional sources with named publishers - the kind of trail a claim needs to survive an editor.
What separates a usable games dataset from a raw dump
The strongest records in this pool share three traits, and each decides whether a project survives contact with production:
- Depth of history: Steam Charts' monthly tables reach each title's own launch, VGChartz' weekly index reaches December 2004, SteamSpy's playtime column accumulates from March 2009, GDELT's event stream from January 1979 and itch.io's archive to its 2013 founding - long enough to fit models on full console generations rather than fragments.
- Identifier hygiene: the numeric appid keys every Steam-side record cleanly against player curves, price history and review corpora; RAWG gives matching pipelines an id, slug, display title and regional alternate names; Speedrun.com hangs every run off stable game, category and level ids. Joins stay clean when identifiers stay stable.
- Declared caveats: SteamSpy publishes owner ranges precisely because false precision would be worse; VGChartz labels every figure an estimate with no guarantee of accuracy; Google Trends values are relative to their comparison set, so two rows can both read 100; and the Pushshift record carries a moderation-only scope. Name the caveats before you build on the numbers.
None of these traits shows up in a screenshot. They show up in sample rows - which is why every Datadory record leads with real rows and a documented field dictionary before any commitment.
Who uses Interactive Home Entertainment data?
Investors and quant researchers build alternative signals: SteamSpy owner bands and monthly player ledgers joined with GDELT event flow around publishers, benchmarked against the 16,598-title sales CSV. Their working rule - carry the band width through the model, and treat week-one figures for new releases as unusable until sampling stabilizes. Their page: interactive home entertainment for investors and quants.
Data scientists and ML engineers train recommendation and ranking systems on RAWG's ratings, tags and similar-game suggestions, then calibrate sampled ownership figures before citing them - anchoring on what is measured (review counts) rather than what is extrapolated (owner midpoints). Their page: interactive home entertainment for data scientists.
Market researchers and consultants size markets by triangulating install-base estimates, weekly retail charts and lifetime platform totals, and profile the indie long tail through 1.5 million listings with price points attached - 220K RAWG-listed developers among them are addressable accounts. Their page: interactive home entertainment for market researchers.
Competitive intelligence and product teams watch whether rival titles compound or decay through concurrency curves and TwitchTracker viewer-hour rankings, and track discount windows across the storefront. Their page: interactive home entertainment for competitive intel teams.
Journalists, academics and students reach for the citable layer: VGChartz archives, adjudicated world-record runs and GDELT's multilingual news coverage back to 1979, when a claim about the games industry has to survive an editor. Their page: interactive home entertainment for journalists and academics.
Delivery: how the games shelf reaches your warehouse
Two translation jobs belong on our side of the line rather than yours. First, storefront-shaped sources arrive built for shoppers, not analysts - turning HTML surfaces and single-purpose feeds into typed columns with stable join keys is work done once upstream so you never repeat it. Second, the estimate-shaped records need interpretation attached: owner ranges kept as ranges, review-derived ratios computed from raw counts, relative interest indexes flagged as relative. Both translations hold across every delivery.
What the numbers say
Every figure below comes straight from the cataloged records, as of August 2026. Against the 7.81 mean across all 1,744 datasets Datadory catalogs, this slice averages 7.75 with ten of sixteen primary records at 8 or higher - a spread carried less by weak sources than by one deliberately low-scored, moderation-scoped outlier. Set the pool against its alternatives in the best interactive-home-entertainment datasets ranking.
Keep reading
Continue into the connected pages:
- interactive-home-entertainment data hub - the full 18-record pool on one index page, with quality scores, field dictionaries, sample rows and coverage spans for every dataset in this guide.
- data scientists - how modeling teams turn ratings corpora, tag vocabularies and playtime aggregates into recommendation and ranking systems.
- market researchers - sizing workflows anchored to install-base estimates, weekly retail charts and the fixed sales benchmark.
| measure | figure |
|---|---|
| Datasets cataloged for the industry | 18 pooled (16 primary + 2 related) |
| Mean quality score, industry slice | 7.75 of 10 (catalog-wide average: 7.81 across 1,744 datasets) |
| Records scoring 8 or higher | 10 of 16 (62.5%), versus 62.8% catalog-wide |
| Records scoring 9 or higher | 5 of 16, topped by Common Crawl Web Corpus at 10 |
| Largest catalogs in the pool | RAWG 500,000+ games and 1.1M ratings; Steam Store 261,285 applications; itch.io 1,509,534 listings |
| Deepest time series | VGChartz weekly index to December 2004; GDELT events to January 1, 1979; SteamSpy playtime since March 2009 |
| Fixed historical benchmark | Video Game Sales with Ratings (Kaggle): 16,598 titles, 1980-2016, four-region unit splits |
| Family | Records | Question it answers |
|---|---|---|
| Live engagement | Steam Web API, Steam Charts Concurrent Players, TwitchTracker Game & Channel Rankings | Is this title growing or fading, right now and over its whole life? |
| Ownership and money | SteamSpy API, VGChartz Video Game Sales Charts, Video Game Sales with Ratings (Kaggle) | How big is the installed base, and how many units moved at retail? |
| Community and competition | Speedrun.com REST API | How fast do players finish it, and how quickly does the record fall? |
| Attention, news and web context | Google Trends, GDELT Project 2.0, Common Crawl Web Corpus, Chrome UX Report, Facebook Graph API, Reddit Data API via Pushshift | What did the surrounding market say, search, read and feel? |
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
How far back does Steam player history go?
To each title's own launch. Steam Charts keeps one row per game per month alongside live standings, 24-hour and all-time peaks - Counter-Strike 2's February 2025 row reads 1,039,662 average players, up 3.60 percent on the month, peaking at 1,818,368. SteamSpy's playtime aggregates accumulate since March 2009, and TwitchTracker holds monthly channel history back years, including the platform's all-time peak month of May 2021.
How accurate are video game sales estimates?
Accurate enough for magnitude, never for invoicing. SteamSpy publishes ownership as ranges such as 100,000,000 .. 200,000,000 precisely because false precision would be worse, and VGChartz labels every retail figure an estimate with no guarantee of accuracy. The calibration habit that works: anchor on what is measured - review counts are observed, not estimated - carry both band bounds through the model, and triangulate against publisher announcements.
Can I track Twitch viewership by game?
Through TwitchTracker - Game & Channel Rankings: average and peak concurrent viewers, hours watched, minutes broadcast, follower gains and share-of-platform percentages per game and per channel. Leaderboards cover trailing seven-day windows, summaries cover thirty days, and per-entity monthly histories stretch back years.
What data can power a game recommendation model?
RAWG is the richest single library: more than 500,000 games across 50 platforms with 1.1M ratings, 58K tags, Metacritic aggregates and machine-generated similar-game suggestions that serve as training ground truth. SteamSpy's tag vectors and average-versus-median playtime pairs add behavioral features, and the gap between a title's mean and median playtime is itself signal about how concentrated its audience is.
Which datasets cover console retail sales?
Two complement each other. VGChartz indexes weekly software charts across Global, USA, Europe, UK, Germany, France and Japan back to December 2004, with lifetime platform hardware totals and a 67,171-title database. The Kaggle Video Game Sales with Ratings CSV fixes 16,598 titles from 1980 through 2016 with North America, Europe, Japan and other-region unit splits as a stable benchmark.
Where do verified speedrun world records come from?
From community adjudication. The Speedrun.com REST API logs hundreds of thousands of runs across tens of thousands of games, each moderated against published rules and carrying a verification status and timestamp beside its performed date. Three timing methods ride side by side down to the millisecond, so comparisons hold only when the clock is held constant - and the gap between performed and verified dates makes moderation latency a measurable quantity.
What does a single Steam store record carry?
One row per application - games, DLC, demos and software - with title, full and short descriptions, developer and publisher, genres, categories and tags, platform availability, system requirements, localized prices with nested discount packages, achievement totals and review-score summaries. Coverage spans 1990s releases still on sale through pre-orders dated years ahead, across 261,285 applications as of August 2026.