Integrated Telecommunication Services · PEGI
PEGI Ratings Database
Datadory delivers pegi ratings database data covering the official Pan European Game Information record for 38 European countries - 42,941 traditionally rated products through end of 2025, each carrying its age label (3, 7, 12, 16 or 18), content descriptors, rating rationale and per-platform release dates back to 2003.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- 38 European countries operating PEGI as their single shared age-rating system - one label valid from Portugal to Poland, unlike territory-split schemes elsewhere
- How far back
- Ratings from 2003 (the search floor) through 2026, with the official statistics series covering products rated through end of 2025; new certificates appear continuously as publishers submit
- How fine
- One record per rated product, with per-platform release dates nested inside each record and rollups available to publisher or age-label level
What is the PEGI Ratings Database?
PEGI Ratings Database is the regulatory-authority anchor of Datadory's integrated telecommunication services data hub - the official record of how Europe rates its games, run by Pan European Game Information under the Interactive Software Federation of Europe. Where World Bank WDI counts mobile subscriptions per 100 people and nPerf crowdsources network speed tests, this corpus answers a different question about the same industry's customers: what is this product allowed to be sold as, and on what stated grounds.
The numbers first. 42,941 products carry a traditional-procedure PEGI rating through end of 2025 - the pre-release examination route, distinct from the IARC questionnaire path used for many digital-only titles. The distribution leans safe: 14,816 records sit at PEGI 3 and another 10,351 at PEGI 12, while only 3,943 reach PEGI 18. Coverage reaches back to 2003, when the search filter floor opens, and runs through 2026 releases. Around the records sits a filterable surface spanning five age labels, fifteen content descriptors, roughly 2,900 publisher entries and every platform family from PS5 and Switch 2 to Google Play and Meta Quest.
What separates a ratings database from a ratings badge is the writing underneath each label. Every record carries the rationale - the actual sentence explaining why a title earned its number - plus a brief outline of the game, content-specific issues naming the scenes that drove the call, and other issues covering mechanics like in-game purchases. That prose is the asset: it makes the corpus machine-filterable and human-explainable at once, which is why it anchors compliance work across the 38 countries that share the system.
What does a sample row look like?
The corpus's most-referenced record and its most-surprising one, rendered as delivered rows:
title : Grand Theft Auto V platform : Xbox Series X|S, PlayStation 5
age_rating : 18 publisher: on request
descriptors: ["violence", "bad_language", "sex", "gambling", "in-game-purchase"]
releases : [{"Xbox Series X|S": "15/03/2022"}, {"PlayStation 5": "15/03/2022"}]
record_id : 109394Read what the shape proves. The descriptors line is the analytical payload: five structured values from a fixed fifteen-term vocabulary, not free-text tags, so 'how often does gambling appear alongside violence' is a group-by rather than a reading project. The releases field keeps per-platform dates inside the rating record - one row per product, dates nested - instead of forcing a join against a separate release table. And record_id gives every row a stable numeric identity, which is what lets a sample cut reconcile cleanly against any prior extract you already hold.
The second row shows the corpus's range in one contrast: Minecraft lands at PEGI 7 with an empty descriptor list and a PC release date of 15/11/2011. Same eleven columns, opposite ends of the severity scale - which is exactly what makes the label distribution a usable feature rather than a constant.
What fields does the dataset include?
Eight fields form the stable spine of every record: identity (title), the classification decision (age_rating, descriptors), and the explanation layer (rating_rationale, brief_outline, content_specific_issues, other_issues, release_dates_platforms). Definitions trace directly to the source's observable pages - result cards and detail views - and hold across the catalog, which is why the same eight columns resolve identically whether a record enters from a title search, a publisher facet or a year drill-down.
Four further fields ride along on record-dependent subsets and fold under additional fields on request: platform strings populate fully on multi-system titles, publisher attribution appears where the source surfaces it beneath the title, the numeric record id is present throughout but quoted case by case since it doubles as the reconciliation key against prior cuts, and localized title variants ship as a joined table across roughly 26 languages rather than as extra columns.
The design point worth noting: this is a shallow-and-clean dictionary, one row per rated product, with the explanation preserved as prose rather than flattened away. Teams that need only labels get a compact table; teams doing content-regulation work keep the sentences that justify each label without leaving the schema.
What does coverage look like across geography, time and granularity?
- Geography - 38 European countries by construction. PEGI is the shared system: one rating issued once is valid across the entire footprint, so there is no per-country reconciliation the way there is between, say, North American and European schemes. For publishers and storefront operators, one query answers the continental question; the ESRB Ratings Search Database covers the US and Canada counterpart jurisdiction with roughly 34,121 certificates since 1994.
- Temporal - 2003 through 2026 on the records themselves, with the official statistics series running through end of 2025. New certificates appear continuously as publishers submit, so the corpus grows by accretion rather than in scheduled batches - and the 2003 floor means two decades of label-mix history for trend work.
- Granularity - one record per rated product, with per-platform release dates nested inside each record and natural rollups to publisher, age-label or year level. Product-level rows answer certification questions; publisher rollups answer portfolio-posture questions.
One honest caveat: the 42,941 count covers traditional-procedure ratings only. IARC-questionnaire titles - largely digital-storefront releases - sit outside it, so the true population of PEGI-labeled products runs higher than the headline. Samples can scope around or through that boundary depending on which population your analysis needs.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
You pick the channel and the cadence; the field dictionary above travels unchanged through all three. Rows arrive flattened to one observation per rated product with the descriptor vocabulary already normalized to the fifteen-term scheme and the explanation prose intact rather than truncated, so your compliance keys line up against ours without a remapping pass. Cadence changes are a settings conversation, not a re-integration project, and a sample cut to your publisher roster comes first either way.
Who uses this data, and for what?
- Competitive intelligence & product teams track competitor titles' PEGI certificates and descriptors as they publish - relevance 3 in their integrated-telecommunication-services pack, the strongest read in the industry - because a rival's rating mix leaks its content strategy: a studio shifting toward PEGI 16/18 output is signaling a portfolio pivot before any press release confirms it.
- Journalists, academics & students cite the official European age-rating record for content-regulation reporting - relevance 3 - since PEGI is the institutional body behind the label, and a citation-grade claim needs the issuer's own record rather than a retailer's badge scrape.
- Data scientists & ML engineers extract age labels and descriptors as training labels - relevance 2 - because a fixed fifteen-term descriptor vocabulary over four decades-adjacent corpus of labeled media is rare supervision that arrives pre-cleaned.
- Market researchers & consultants analyze rating mix across the 38-country footprint and years 2003-2026 - relevance 2 - reading severity distributions as a proxy for how regional content appetites and regulation interact.
- E-commerce operators fill PEGI age-rating fields on EU storefront listings - relevance 2 - replacing hand-keyed compliance columns with a sourced record keyed on the official id.
- Investors & quant researchers check publisher exposure to adult-rated titles in European markets - relevance 1 - a diligence screen that reads concentration risk out of a publisher's label distribution.
Which personas get the most value?
Competitive-intel teams and journalists/academics both hold this dataset at relevance 3 - the only double-three tie among the eight datasets in the integrated-telecommunication-services pool, and for complementary reasons. The intel teams want the leading indicator: certificates publish ahead of launches, so descriptor drift on a rival's slate reads as strategy. The journalists want the trailing authority: when a regulator or a parent group asks what a game is rated and why, the answer has to come from the issuing body's own record, with the rationale sentence attached. Data scientists, market researchers and e-commerce operators hold relevance 2, each consuming a different column of the same rows - training labels, severity time series, listing-compliance fills respectively. Developers & builders and investors/quants hold relevance 1, using it as a reference layer rather than a primary corpus. The constant across all six: the label-plus-descriptors pair. Everything here compounds when joined against the pool's reception scores from Metacritic Games and catalog metadata from the IGDB API (Twitch/Amazon), turning ratings into a feature dimension rather than a standalone fact.
What should I know before requesting a sample?
Four things, all knowable upfront. First, population honesty: the 42,941 headline covers traditional-procedure ratings only; IARC-questionnaire titles are real but uncounted publicly, so decide which population your work needs and the sample will be cut accordingly. Second, sparsity is structural: publisher attribution and multi-platform strings populate where the source surfaces them, which is why those columns fold under additional fields on request rather than pretending every record is fully dressed. Third, the prose columns - rationale, outline, content-specific issues - vary in length by orders of magnitude across records; they are stored complete, but plan storage and token budgets for the heavy tail rather than the median. Fourth, scoping is the default: samples cut to your publisher roster, platform set or year window arrive before any recurring delivery is configured, and the underlying content carries its own terms at the source level while Datadory delivers normalized rows under ours - the distinction is spelled out in the sample paperwork before anything ships.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
title | string | Name of the rated game as it appears on the result card and detail page; the human lookup key. | Minecraft |
age_rating | integer | PEGI age label: 3, 7, 12, 16 or 18, with Parental Guidance markings on some digital titles. | 18 |
descriptors | text | Content descriptors attached to the record from the fixed fifteen-term vocabulary. | ["violence", "bad_language", "gambling"] |
rating_rationale | text | Narrative explanation of why the age label was assigned, published verbatim on the detail page. | "This game has received a PEGI 18 which restricts availability to ADULTS ONLY..." |
brief_outline | text | Publisher-supplied synopsis of the game under 'Brief outline of the game'. | on request |
content_specific_issues | text | Detailed description of the specific content that drove the rating decision. | on request |
other_issues | text | Additional notes outside the descriptor set, such as in-game purchase mechanics or online interaction. | on request |
release_dates_platforms | text | Per-platform release dates listed on the detail page in DD/MM/YYYY form. | PC - 15/11/2011 |
Questions buyers ask
How many games does the PEGI Ratings Database cover?
42,941 products rated under the traditional pre-release procedure through end of 2025, per PEGI's own statistics: 14,816 at PEGI 3, 10,351 at 12, 7,381 at 7, 6,360 at 16 and 3,943 at 18. Titles rated through IARC questionnaires sit outside that count, so the full labeled population runs larger than the headline figure.
Which countries recognize the PEGI label?
38 European countries operate PEGI as their single shared age-rating system. One certificate issued once is valid across the entire footprint, so a compliance check against the corpus answers for a whole continent at once - the property that makes it cheaper to work with than territory-by-territory schemes.
What content descriptors does the dataset track?
Fifteen: violence, fear, bad language, sex, gambling, drugs, discrimination, in-game purchases, horror, pressure to play, safe online gameplay, time-limited offers, cryptocurrency, unrestricted communication and paid random items. The digital-economy marks make the vocabulary useful for monetization research, not only content-severity studies.
Does each record explain why the rating was assigned?
Yes. Records carry the written rating rationale, a brief outline of the game, content-specific issues naming what drove the decision, and other issues such as in-game purchase mechanics. The explanation layer is what turns an age label into an auditable compliance artifact rather than an opaque number.
Can I get PEGI ratings joined to critic scores?
Within Datadory's catalog, yes by construction: the Metacritic Games corpus holds 14,330 scored titles in the same industry pack, and the IGDB metadata layer carries board age ratings across its relational entities. A sample can demonstrate the join on your title list before any recurring configuration.
Can the sample be scoped to specific publishers or years?
Name the publishers, platforms, years or age labels and the sample arrives shaped to that scope with the complete field dictionary attached. Samples precede any commitment, and the schema validated in the sample is the schema the recurring feed keeps.
Notes on this record
- One label, 38 countries A PEGI certificate issued once is valid across every country operating the system, so continental compliance checks resolve in one query instead of 38 lookups.
- Prose, not just numbers Rationale, outline and content-specific issues ride on every record, making each label auditable - rare among rating corpora, which usually stop at the icon.
- A fixed 15-word vocabulary Descriptors normalize to one closed term set spanning violence to paid random items, so severity mix and monetization marks both become group-bys rather than parsing projects.
- Dates inside the record Per-platform release dates nest within each rating row, keeping one row per product instead of splitting releases into a second table to re-join.
- Stable numeric ids Every record carries its PEGI record id, giving samples and recurring feeds a shared reconciliation key across cuts delivered months apart.
- Where it sits in the slice Quality score 7 among eight cataloged Integrated Telecommunication Services sources - the regulatory-authority counterweight to World Bank penetration statistics, nPerf network maps and reception aggregators.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.