Datadory notebook
RAWG API Similar Games Endpoint: Look-Alike Data, Delivered
Datadory delivers RAWG similar games data covering more than 500,000 games across 50 platforms - machine-ranked look-alike pairs generated by computer vision applied to game imagery, shipped as an edge list keyed by game ID, with Metacritic aggregates, 0-5 community ratings, average playtimes and age ratings riding on every row. Name the franchises or genres you care about; real rows come back before you commit.
1,744 datasets. Pick your catch.
What the RAWG similar-games data actually is
What the market searches for as RAWG similar games data, Datadory delivers as rows: a look-alike edge list where every entry runs from one game ID to the titles a model judges similar to it. The pairs are machine-ranked from computer vision applied to game imagery, not assembled from overlapping tag lists - which is exactly why they behave differently from what a genre filter returns. Look-alikes cross studios, publishers and storefronts, because the model reads what the games look and feel like rather than what shelf label somebody filed them under.
The edge list does not travel alone. It sits inside the widest supply-side table in the Interactive Home Entertainment pool: RAWG Video Games Database API data, more than 500,000 games across 50 platforms, ringed by roughly 2.1 million screenshots, 1.1 million ratings, 58,000 tags and 45,000 publishers. Each title carries its own count of attached suggestions, so you can weigh how strongly the model committed to every seed game. Child tables for DLC additions, development-team credits, game-series membership and parent-franchise links arrive alongside, which means algorithmic neighbours and lineage neighbours stay separable instead of blurring into one undifferentiated "related" pile.
On quality, this slice scores 9 out of 10 in the catalog against a 7.81 average across all 1,744 datasets, carried by unusually disciplined field documentation - the verified confidence flag on field definitions is shared by 1,495 records, some 85.7 percent of the catalog.
What travels with each look-alike row
An edge list on its own answers "what resembles this." The columns around it answer "why that matters," and they ride on the same rows:
- Identity: numeric game ID, slug, display title, original store title and regional alternates - several independent chances for a matching pipeline to survive re-releases and localisations.
- Release facts: the release date with a to-be-announced flag that keeps the forward calendar queryable, plus per-platform dates.
- Reception: the Metacritic aggregate beside a 0-5 community rating with its bucket distribution and count - two verdicts, one row.
- Behaviour: average playtime in hours and the library breakdown - who marked the game owned, playing, beaten or to-play.
- Governance and reach: the age rating and the platform set, drawn from 50 platforms including mobile.
The shape looks like this:
# Similar-games row -- shape preview; value slots fill in your sample
game_id : 3498
game_name : Grand Theft Auto V
similar_id : <look-alike game id>
similar_name : <look-alike display title>
metacritic : <critic aggregate> rating : <community score, 0-5>
playtime : <average hours> added : <owned/playing/beaten/to-play>
released : <release date> platforms : <platform set>
esrb_rating : <age rating>
suggestions_count : <look-alikes attached to the seed title>One lean to model around rather than discover late: store links and age ratings skew toward US-facing metadata, so non-US work wants that tilt handled explicitly.
Six products behind a production look-alike feature
Similarity answers "what should we show next." A feature that survives contact with players needs three more answers - is anyone actually there, how big is the base, and what did it sell - and each lives in its own product on the same shelf.
RAWG Video Games Database API generates candidates: vision-derived look-alike pairs across all 50 platforms, with scores, playtimes and age ratings on every row.
Steam Web API data confirms the present tense - point-in-time concurrent players and global achievement unlock percentages for the PC slice, read at capture. Counter-Strike 2 stood at 1,001,216 concurrent players when the feed was verified during August 2026 cataloging.
Steam Charts - Concurrent Players holds the only month-by-month average-player ledger in the pool, one row per game per month back to each title's launch, with absolute and percentage gains and monthly peaks.
SteamSpy API contributes the ownership side - estimated owner ranges plus average and median playtime accumulated since March 2009, published as estimates and cited as estimates.
itch.io Indie Games Browse stretches candidates past the charted top few hundred indie releases a year into 1,509,534 tagged listings as of August 2026, prototypes and jam entries included.
Video Game Sales with Ratings anchors the money claim: 16,598 titles that each cleared 100,000 copies, one flat schema held stable since 1980, regional unit sales split North America, Europe, Japan and elsewhere.
How the pieces assemble into a ranked row
The stack most teams converge on has four moves, and the order matters.
Run end to end, the pattern is simple to state: look-alike edges supply candidates, Steam-side products supply behavioural weight, and the sales benchmark supplies the economic floor. Every layer speaks game ID, so assembly is joins, not glue.
Cross-platform breadth against PC depth
The nearest peer to the RAWG catalog is Valve's telemetry, and the split is breadth versus behaviour. RAWG answers similarity across 50 platforms with Metacritic aggregates, age ratings and vision-derived look-alikes; Steam reports live concurrency, achievement unlock rates and, with user consent, owned libraries - but only for Steam titles, and neither critic scores nor review aggregates appear anywhere in it.
A practical reading: the cross-platform candidate generator comes from the broad catalog, the engagement ranking comes from Steam-side products, and the two meet on game IDs. Where they disagree - a title the model loves that players have quietly abandoned - is usually where the interesting question lives. The head-to-head RAWG vs Steam Web API comparison settles the catalog-versus-telemetry question cell by cell.
Who builds on look-alike data
- Data Scientists & ML Engineers - half a million titled rows with dual score systems, playtimes and ready-made look-alike edges: evaluation truth for a similarity function and feature material for success-prediction models.
- Market Researchers & Consultants - genre and franchise sizing across PC, console and mobile from one consistent catalog instead of per-storefront fragments.
- Competitive Intelligence & Product Teams - rival catalogs' breadth, scoring and platform spread as a continuous signal rather than a quarterly scramble.
- Developers & Data-Product Builders - discovery surfaces, storefront widgets and catalog search starting from measured pairs rather than shared labels.
- Investors & Quants - catalog-scale supply metrics for diligence on publishers and platforms between filings.
- Journalists, Academics & Students - citable scale figures and per-title reception history for work on how the games market is actually structured.
How the feed is delivered
Name the seed titles and the fields - real look-alike rows come back before you commit.
A scoped sample pulls actual pairs for the games you name, with the field dictionary and coverage notes attached, so validation happens on your data rather than on a screenshot.
The plumbing was always the expensive part of recommendation work. It ships done.
Where to go next
Start with the Interactive Home Entertainment data guide for the whole pool, or skim the best interactive home entertainment datasets list for the ranked view. Row-level, request a sample from RAWG Video Games Database API for the candidates and Steam Charts - Concurrent Players for the engagement weighting. Everything in the slice sits grouped on the interactive home entertainment data hub.
Adjacent questions get their own posts: how do I get Steam concurrent player counts covers the behavioural half of the join, and the RAWG data source profile introduces the publisher behind the catalog.
| Product | Similarity signal | Scale | Best used for |
|---|---|---|---|
| Steam Charts - Concurrent Players | Month-by-month average-player ledger back to each title's launch | Tens of thousands of tracked titles | Post-launch decay curves behind ranking rules |
| itch.io Indie Games Browse | Tags, price and platform coverage across the indie long tail | 1,509,534 listings as of August 2026 | Candidates mainstream charts never reach |
| Video Game Sales with Ratings | Fixed genre and publisher benchmark with regional unit sales | 16,598 titles above 100,000 copies, 1980-2016 | Anchoring the revenue claim attached to a slot |
| Question you are asking | Product to pull | Figure to reach for |
|---|---|---|
| What should we show next to this player? | RAWG Video Games Database API | Machine-ranked similar-game edges |
| Is the look-alike actually alive right now? | Steam Web API data | Concurrent players at capture |
| Did the recommended title hold its audience after launch? | Steam Charts - Concurrent Players | Monthly average players with percentage gain |
| How big is the installed base behind the suggestion? | SteamSpy API | Owner ranges |
| What did comparable titles sell, by region? | Video Game Sales with Ratings | Regional unit sales per title-platform |
| Are we ignoring the long tail? | itch.io Indie Games Browse | Tagged listings beyond the charted top |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
RAWG Video Games Database API
slug · name · name_original …+24 more
Steam Charts - Concurrent Players
appid
SteamSpy API
appid · owners · average_forever …+2 more
Video Game Sales with Ratings
Rank · Global_Sales · Platform …+3 more
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What is the RAWG similar games dataset?
A machine-ranked look-alike edge list covering more than 500,000 games across 50 platforms: every entry points from one game to the titles a model judges similar to it, delivered keyed by game ID with reception and behaviour columns attached to both ends of the pair.
How are the similar-game pairs generated?
By computer vision applied to game imagery rather than by overlapping tag lists, which is why the look-alikes cross studios and storefronts instead of stopping at franchise neighbours. Each title also carries a count of the suggestions attached to it, so the strength of the model's opinion is measurable per game.
Does similar-games coverage include mobile and retro titles?
Yes. The catalog spans more than 500,000 games across 50 platforms including mobile, reaching from early arcade and retro eras through announced-but-unreleased titles still flagged to-be-announced. One lean worth modeling around: store links and age ratings skew toward US-facing metadata.
Can the look-alike feed power a feature inside my own product?
Yes. Delivered rows land in your pipeline already typed and joined, so rankings render inside your product rather than as cached copies of somebody else's lists. Enrichment columns such as scores, playtimes and age ratings ride along on the same rows.
Which other datasets complete a recommendation feature?
Steam Web API data confirms live demand, Steam Charts holds the month-by-month engagement ledger, SteamSpy contributes ownership and playtime estimates, itch.io extends candidates into the indie long tail, and Video Game Sales with Ratings anchors any revenue claim attached to a slot.