Interactive Home Entertainment Data: Player Counts, Catalogs and Sales · Head-to-head

GDELT Project 2.0 vs Video Game Sales with Ratings (Kaggle)

Which interactive home entertainment data: player counts, catalogs and sales data fits your job: GDELT Project 2.0, or Video Game Sales with Ratings. API, files, or your warehouse. Daily, weekly, or hourly.

Interactive Home Entertainment Data: Player Counts, Catalogs and Sales Nearly every country · January 1

GDELT Project 2.0

Interactive Home Entertainment Data: Player Counts, Catalogs and Sales Four regions: North America · Titles released 1980 through 2016

Video Game Sales with Ratings (Kaggle)

Where the fields line up

1 shared field — join on these.

Field GDELT Project 2.0 Video Game Sales with Ratings
Year Year of the event in YYYY form. Year of the game's release

Coverage, side by side

GDELT Project 2.0 Video Game Sales with Ratings
Geographic Nearly every country, georeferenced to city and landmark centroids Four regions: North America, Europe, Japan, rest of world
Temporal January 1, 1979 to present; daily historically, timestamps resolved to fifteen minutes from February 2015 onward Titles released 1980 through 2016, frozen at the October 2016 capture
Granularity One row per event per source document, plus per-mention and per-codified-document layers One row per game title per platform, at annual release-year grain

What each contains

Pick by fit, not by loyalty.

GDELT Project 2.0 Video Game Sales with Ratings
Steward GDELT Project (Google Jigsaw) Community upload on Kaggle
Subject lens Worldwide news codified into events, mentions and a knowledge graph, scored for conflict-cooperation and tone Console game sales snapshot by title, platform, genre, publisher and region
Field dictionary 21 documented fields, verified during research 11 documented fields, verified during research
Ratings content None - tone and Goldstein scores stand in for judgment None, despite the name
Geographic coverage Nearly every country, georeferenced to city and landmark centroids Four regions: North America, Europe, Japan, rest of world
Temporal coverage January 1, 1979 to present; daily historically, timestamps resolved to fifteen minutes from February 2015 onward Titles released 1980 through 2016, frozen at the October 2016 capture
Granularity One row per event per source document, plus per-mention and per-codified-document layers One row per game title per platform, at annual release-year grain
Scale Quarter-billion-plus event records; about 2.5 TB uncompressed per year of knowledge graph 16,598 rows in one table, about 1.4 MB
Best for Media narrative: who was covered, where, how intensely and in what tone Market structure: which games sold, where, on which platforms

What each does better

GDELT Project 2.0

Density per record, proven with a row. One event reads: GlobalEventID 1319329086, Actor1Code GBR (UNITED KINGDOM) meeting Actor2Code BUS in Cambridge, Cambridgeshire, United Kingdom, CAMEO action 051, GoldsteinScale 3.4, carried by 4 articles, average tone +3.65, stamped 20260821140000. The documented dictionary captures 21 of the format's nearly 60 attributes per event, so most of that specification survives any normalization.

Judgment machinery. Every event lands on a -10 to +10 conflict-cooperation scale and a -100 to +100 tone scale, and the knowledge-graph layer runs GCAM across 2,300+ emotions and themes per document. No sales table grades its rows.

Geography to the ground. Actors and actions alike georeference to city and landmark centroids across nearly every country, drawn from hundreds of thousands of outlets with deliberate emphasis on non-Western local media - and machine translation folds 65 languages representing 98.4% of daily non-English volume into the same schema.

Depth of archive. Events run back to January 1, 1979 at daily resolution, with special collections adding 215 years of digitized books, 50+ years of human rights material and closed captioning from 100+ US TV stations.

Video Game Sales with Ratings

Legibility. Eleven columns, one table, 16,598 rows. A reader learns the whole dictionary in minutes and needs no codebook to interpret a cell - against a CAMEO hierarchy, a Goldstein scale and a six-dimension tone vector on the other side.

Commercial ground truth. The figures are unit sales, not sentiment: Wii Sports at 41.49M in North America, 29.02M in Europe, 3.77M in Japan, 8.46M elsewhere, 82.74M globally. No volume of news coverage produces that number; purchasing did.

Benchmark gravity. Four decades of releases in a stable, self-contained shape have made it a default proving ground for forecasting and coursework - 2,027 linked notebooks were attached to it as of August 2026, and its Name, Platform, Genre and Publisher columns join cleanly on plain strings.

Regional honesty. The North America / Europe / Japan / other split is the exact cut a go-to-market team argues about, delivered on every row rather than requiring aggregation.

Where they're equivalent

More than their vocabularies suggest. Both dictionaries were verified during research. Both are flat, one-row-per-thing tables: one row per event per source document on the GDELT side, one row per game per platform on the sales side. Both carry a literal Year column. Both compress the planet into bounded geography - city-centroid coordinates on one side, four sales buckets on the other. And both share a defining absence: neither dataset contains a rating. GDELT grades events for conflict-cooperation and tone, but publishes no stars; the Kaggle file is named "with Ratings" and ships eleven sales columns with no critic or user score anywhere - see user ratings and review counts for what genuine rating fields look like.

They also share their limits, asymmetrically. The sales table is a static snapshot, captured once and never extended past its October 2016 capture, so any post-2016 release is invisible. GDELT keeps moving but retains no per-subject ledger either: it knows an event happened and how the press took it, not how a franchise performed cumulatively. Neither answers a question about money directly - one counts mentions, the other counts units.

The verdict

Verdict: sample both, pick by fit - they are instruments pointed at entirely different questions.

Take GDELT Project 2.0 if your question is about narrative. How the press treated a launch or a publisher, where coverage clustered, whether tone shifted after an incident, which countries owned a storyline, how salient an event was in mentions and articles. Accept its frame: you are reading codified news, so volume signals attention rather than purchase intent, and no revenue appears anywhere.

Take Video Game Sales with Ratings (Kaggle) if your question is about the market itself. Which titles sold, on which platforms, in which regions, across console generations from 1980 to 2016 - and who published them. Accept its frame: it stops in 2016, it counts units rather than dollars, and despite the name it contains no ratings whatsoever.

Sample both, pick by fit. See GDELT Project 2.0 · See Video Game Sales with Ratings

Or take both in one feed

Yes - and the merge is where the pairing starts earning its keep: attach news narrative to sales reality. The sales table supplies the timeline of what shipped and what it moved in units; GDELT supplies the coverage curve around it - when a franchise or publisher entered the world's press, in which countries, with what tone, at what mention volume.

Three alignments decide whether the merge holds. First, entity: CAMEO organization codes and plain-text publisher strings never meet directly, so alignment runs through normalized publisher, franchise and title names matched against event actor names and document themes. Second, place: city-level geocoding has to roll up to the same four regions the sales table uses - North America, Europe, Japan, rest of world - which trades away some precision on both edges. Third, time: a release year must be joined to the window of coverage surrounding it, since the sales table carries no date finer than a year while GDELT resolves to fifteen minutes. We handle all three in the merge, delivered daily, weekly, or hourly - your call. Or take both in one feed.

API, files, or your warehouse. Daily, weekly, or hourly.

Fair questions

Is GDELT Project 2.0 better than Video Game Sales with Ratings (Kaggle)?

Better at different jobs. GDELT owns narrative: 21 documented fields per codified event - CAMEO actor pairs, Goldstein conflict-cooperation scores, tone, city-level georeferencing and media-salience counts - drawn from news in over 100 languages with history reaching back to 1979. The Kaggle table owns the market: 11 documented fields over 16,598 games with regional unit sales by platform, genre and publisher, covering releases from 1980 through 2016. Sample both, pick by fit.

Do the two datasets cover the same ground?

Only structurally. Both are one-row-per-entity tables with a date axis, named parties, a classification scheme and numeric magnitudes - even a column called Year on both sides. Substantively they barely touch: GDELT records what the world's press said, where and in what tone; the sales table records how many copies of each game sold in each region. Their one substantive meeting point is geography, and even there GDELT georeferences to cities while the sales table stops at four regions.

Does Video Game Sales with Ratings actually contain ratings?

No. Despite the name, the file as served carries eleven sales columns - rank, name, platform, year, genre, publisher and five sales figures in millions of units - with no critic or user scores anywhere in it. The nearest thing to a grade in this pairing sits on the GDELT side, which scores every event for conflict-cooperation from -10 to +10 and tone from -100 to +100. Anyone whose brief requires star ratings needs a different instrument.

Which dataset should a market analysis sample first?

The sales table, to establish what happened commercially: platform cycles, genre volumes, regional skews and publisher dominance across four decades of releases, readable in minutes. Then bring in GDELT for the narrative layer - how press attention and tone moved around the titles and companies that surface on the sales board. Both arrive normalized to their documented field dictionaries.

Can Datadory deliver both datasets together?

Yes - alone or merged onto one timeline, delivered daily, weekly, or hourly, your call. Each arrives normalized to its documented dictionary (21 fields on the GDELT side, 11 on the sales side) with sample rows for inspection before anything ships. The joining work is name-based alignment between CAMEO actor codes and publisher strings, rolling city-level geocoding up to four sales regions, and pinning event timestamps to release years. Or take both in one feed.