Interactive Home Entertainment Data: Player Counts, Catalogs and Sales · Head-to-head
itch.io - Indie Games Browse vs Steam Store - Scrapeable Catalog
Which interactive home entertainment data: player counts, catalogs and sales data fits your job: itch.io - Indie Games Browse, or Steam Store - Scrapeable Catalog. API, files, or your warehouse. Daily, weekly, or hourly.
itch.io - Indie Games Browse
Steam Store - Scrapeable Catalog
Where the fields line up
1 shared field — join on these.
| Field | itch.io - Indie Games Browse | Steam Store - Scrapeable Catalog |
|---|---|---|
platforms | Delivery targets flagged on the release: windows, mac, linux, android or browser-playable HTML. | Booleans for windows, mac and linux availability. |
Coverage, side by side
| itch.io - Indie Games Browse | Steam Store - Scrapeable Catalog | |
|---|---|---|
| Geographic | One global marketplace; USD-denominated default pricing with no regional partition | Global storefront; the same title prices differently per country context |
| Temporal | Live current state; publication timestamps per listing and an archive reaching to the 2013 launch | Live current state; release dates spanning 1990s re-releases through future-dated pre-orders |
What each contains
They tie on 1 attribute. Pick by fit, not by loyalty.
| itch.io - Indie Games Browse | Steam Store - Scrapeable Catalog | |
|---|---|---|
| Operator | itch corp | Valve |
| Shelf described | Independent and jam games published on itch.io, with game assets, tools, soundtracks and physical games in the wider catalog | PC storefront applications: games, DLC, demos and software distributed on Steam |
| Subject lens | Live per-listing facts: title, creator, blurb, genre, price string, currency, platform targets, publication timestamp, community tags, cover art | Live per-application spec: identity, price tiers, discounts, descriptions, languages, requirements, genres, review summaries, achievements, release dates |
| Field dictionary | 12 documented fields, verified during research | 30 documented fields, verified during research |
| Scale | 1,509,534 browse results; about 1.47 million catalog results as of August 2026 | 261,285 searchable applications observed |
| Geographic coverage | One global marketplace; USD-denominated default pricing with no regional partition | Global storefront; the same title prices differently per country context |
| Temporal coverage | Live current state; publication timestamps per listing and an archive reaching to the 2013 launch | Live current state; release dates spanning 1990s re-releases through future-dated pre-orders |
| Audience signal | None carried - no rating, review or ownership field exists in the dictionary | Review-summary label with positive/negative splits, recommendation counts, recent-review totals, Metacritic where available |
| Creator dimension | Author name anchoring a creator page that holds the studio's entire upload history | Developer and publisher name lists inside each application record; no creator-page spine |
| Platform scope | windows, mac, linux, android and html (browser-playable) | windows, mac and linux booleans |
| Best for | Long-tail indie supply, jam intelligence, pricing-floor research and creator-level tracking | Commerce depth, sentiment signals, engineering specifications and lifecycle staging |
What each does better
itch.io - Indie Games Browse
Long-tail breadth nothing else captures. 1,509,534 browse results against a search index of 261,285 - and the gap is not uniform, it concentrates exactly where mainstream charts go blind. A two-person horror prototype uploaded for a jam weekend exists here as a fully attributed row and nowhere in the Steam ledger at all.
The creator dimension. Every listing carries an author name that anchors to a creator page holding that studio's complete upload history, so output-per-studio and follow-the-developer studies run off one column. Steam's dictionary knows developers only as a name list inside each app record.
Publication timing at listing resolution. createDate stamps when each entry appeared - one sample reads Fri, 21 Aug 2026 13:47:40 GMT - so new-supply series can be built at whatever grain the question needs, and jam weekends surface as visible spikes across roughly thirteen years of archive reaching to the platform's 2013 launch.
Platform honesty. windows, mac, linux, android and html in one enum - a browser-playable entry like Serperwave, the vaporwave Snake reimagining in the sample rows, is flagged as playable in a tab, a fact the desktop-only Steam booleans cannot represent.
Steam Store - Scrapeable Catalog
Depth per title, proven with a row. One live record reads: type game, name Dota 2, steam_appid 570, is_free true, developers [Valve], publishers [Valve], platforms true across windows, mac and linux. Every application arrives that completely specified, and the full payload extends to descriptions, language support, requirements, genres, screenshots and movies.
Audience sentiment the other side cannot see. Dota 2's recent-review panel reads Very Positive with 5,430 positive against 689 negative out of 6,119 reviews; recommendations totals positive votes, review_score_desc grades each title on Steam's familiar label ladder, metacritic appears where scored, and total_reviews counts the recent window - see user ratings and review counts. The itch.io record carries no rating, review or engagement field whatsoever.
Engineering and governance attributes. Minimum and recommended requirements per operating system, supported_languages down to subtitle depth, achievement totals, required_age gates and board-rating descriptors make compatibility audits and audience-suitability screens a filter rather than a research project - see content rating fields.
Lifecycle span. release_date pairs a coming-soon flag with dated releases reaching from 1990s back-catalog reissues to pre-orders years ahead, so the ledger covers both history and pipeline.
Where they're equivalent
More than the field counts suggest. Both dictionaries were verified during research, and both records score 8 out of 10. Both are first-party catalogs - each operator describing its own shelf - so attribute definitions inherit the store's own vocabulary rather than a third-party compiler's.
They also share their limits, symmetrically. And neither exposes individual reviewer identities: Steam aggregates its reviews into scores and counts, itch.io omits reception entirely, so reviewer-level analysis needs a different instrument on either side.
The verdict
Verdict: sample both, pick by fit - they measure different halves of one industry.
Accept its frame: no sentiment, no ownership, desktop-centric questions underserved.
Accept its frame: the long tail barely registers, and browser or mobile targets are absent.
Market researchers usually start on the itch.io side for denominators and move to Steam for value; investors and quants tend to invert that order; journalists and academics studying jam culture have no substitute for the census at all.
Sample both, pick by fit. See itch.io - Indie Games Browse · See Steam Store - Scrapeable Catalog
Or take both in one feed
Yes - they stack into one games picture that neither completes alone. Then enrich the matched titles on the Steam ledger: purchase tiers and discounts, review-summary labels, requirement footprints, language coverage, lifecycle stage.
Three seams decide whether the merge holds. First, identity: numeric data-game_id and steam_appid come from unrelated id spaces, so alignment runs through normalized title-plus-creator strings - see app identifier fields. Third, platform vocabulary: android and html flags on one side meet windows-mac-linux booleans on the other, needing a small mapping layer before reach comparisons mean anything. We handle all three in the merge. Browse the rest of the shelf at the interactive home entertainment data hub.
Datadory ships either record alone or both merged onto one title list, delivered daily, weekly, or hourly - your call. Or take both in one feed.
API, files, or your warehouse. Daily, weekly, or hourly.
Fair questions
Do the two datasets cover the same ground?
Six concepts deep. Everything else diverges by design: itch.io spends its columns on creators, publication timing and a thirteen-year indie archive; Steam spends its extra eighteen fields on purchase tiers, review summaries, system requirements, languages and lifecycle. One is a census of what got made, the other a ledger of what it sells for and how it is received.
Which dataset documents more fields per title?
Steam's, by eighteen: 30 documented fields against itch.io's 12. Density versus reach, in schema form.
Which dataset should an indie-market sizing project sample first?
Both land normalized to their documented field dictionaries.
Can Datadory deliver both datasets together?
Yes - alone or merged onto one title list, delivered daily, weekly, or hourly, your call. Each arrives normalized to its documented dictionary (12 fields on the itch.io side, 30 on the Steam side) with sample rows for inspection before anything ships. The joining work is title-and-creator alignment between unrelated numeric identifiers, reconciling pricing-floor strings against tiered cents ladders, and mapping android-and-browser platform flags onto desktop booleans, which we handle in the merge. Or take both in one feed.