Google Data Commons
Datadory delivers industrial conglomerates data covering Google Data Commons: tens of thousands of harmonized statistical variables - manufacturing value added, GDP, employment, trade and population - drawn from the World Bank, OECD, Eurostat, FAO, WHO, NOAA, EPA, BLS, BEA and the U.S. Census Bureau, resolved place by place from worldwide down to county level with every figure attributed to its source facet, delivered as typed, joinable rows.
What is Google Data Commons?
The shortest honest description of Google Data Commons is also the most useful one: somebody took the world's public statistics and gave them a single shape. Run as a Google initiative, it harmonizes figures from the World Bank, OECD, Eurostat, the Food and Agriculture Organization, the World Health Organization, NOAA, the US Environmental Protection Agency, the Bureau of Labor Statistics, the Bureau of Economic Analysis and the U.S. Census Bureau - among hundreds of provenance sources - into one statistical knowledge graph. Every indicator becomes a statistical variable tied to an entity and a date: a variable like Manufacturing_ValueAdded_AsFractionOfGDP resolves against country/USA rather than living in one agency's spreadsheet. The consequence for anyone tracking diversified industrials is that manufacturing value added, GDP, employment, trade and population arrive for every market a conglomerate operates in, in one consistent form, instead of a dozen agency portals each speaking its own dialect.
Inside Datadory's catalog this record anchors the macro layer of Industrial Conglomerates - the measure-the-pond layer that sits beneath company filings and facility registries. Get a sample of this dataset and we will cut it to the variables, countries and years you name.
What do sample rows look like?
Three shipped observations, flattened for reading:
# one observation = one entity x one variable x one date x one provenance facet
variable : Count_Person
entity : country/USA
date : 2020
value : 329484123
variable : Count_Person
entity : country/DEU
date : 2020
value : 83160871
facetId : 18369491376878146239
entity : country/DEU
date : 2020
value : 83160871The third row is the one worth staring at. It repeats Germany's 2020 population but tags it with a facetId - a provenance stamp identifying which source's measurement is being reported. When two agencies publish different counts of the same thing, the graph does not average the disagreement away; it keeps both, facet by facet, so a pipeline chooses its authority deliberately instead of discovering the conflict downstream. In your sample each observation lands as a typed row - variable, entity, date, value, facet - ready to filter by region, indicator or source.
What fields does the dataset include?
Ten documented fields make up an observation response, grouped four ways. An addressing layer (byVariable, byEntity) keys results by statistical variable and by place, so one pass fans out across both dimensions without repeated lookups. A provenance layer (orderedFacets, facetId) carries the source attribution behind every number - the feature that separates this from an ordinary indicator table. A measurement layer (observations, date, value) holds the figures themselves. A plumbing layer (obsCount, earliestDate/latestDate, nextToken) reports how much returned, how far history runs and where pagination continues.
Five of the ten carry unambiguous presentation and ship with worked examples in the dictionary below. The remainder - the facet list, the observation arrays, the count and bounds fields, the continuation token - arrive populated on records but without a fixed worked example, so they fold under additional fields on request, together with the derived columns Datadory attaches on delivery: cross-facet reconciliations, place-hierarchy rollups and long-run backfills.
What does coverage look like across geography, time and granularity?
Geography - worldwide at the top, descending to country, state and district or county level. The United States runs deepest of all: state and county grain that most international collections never attempt, which matters when a conglomerate's footprint is measured plant by plant.
Temporal - declared per series rather than asserted site-wide. Multi-decade runs are common, census and World Bank lines reach back to the 1960s, and every provenance facet self-reports its own earliest and latest dates. Available history is therefore something you read off the data, not something a brochure promises.
Granularity - one observation per entity per statistical variable per date per provenance facet. That fourth dimension is the discipline here: aggregation always starts from the same atomic unit, so nothing is silently blended across disagreeing sources.
Against the wider catalog - 1,744 datasets, average quality score 7.81 - this slice scores 7/10, carried by verified field documentation and shipped sample rows, tempered by a scale the source describes in charts rather than published counts.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Pick the channel your team already works in: queryable rows for live lookups on specific variables and markets, flat files sized for overnight warehouse loads, or a direct pipe into Snowflake, BigQuery or Redshift. Cadence is yours to set - and to change when the models change.
Every delivery ships the full field dictionary, sample rows for validation, and a schema that holds steady between deliveries.
Who uses this data, and for what?
- Multi-market macro baselines. Pull manufacturing value added, GDP, employment and trade indicators for every country a conglomerate touches, from one consistent shape - the denominator work that feeds the wider market sizing shelf.
- Provenance-aware benchmarking. Facet tagging makes source disagreements visible instead of hidden, so a Germany-versus-France comparison can name the authority behind each figure.
- Long-run trend construction. Series reaching back to the 1960s turn a snapshot into a structural story - deindustrialization, catch-up growth, energy transitions - with dates declared per facet.
- Feature engineering for models. Typed entity-variable-date rows join cleanly onto firm-level feeds; see the ML model training workflows built on exactly this shape.
- Citation-grade research. Every figure carries its source, which is why the citation-first teams start here; the citation-grade research playbook leans on it.
Which personas get the most value?
Ranked by relevance in Datadory's persona tagging. Data scientists and ML engineers (relevance 3) get harmonized macro features joined on stable entity keys - see the data scientists x industrial conglomerates workflows. Market researchers and consultants (relevance 3) replace a shelf of agency portals with one queryable shape - market researchers use cases. Developers and data-product builders (relevance 3) build on an addressing scheme that treats places and variables as first-class keys - developers builders use cases. Journalists, academics and students (relevance 3) cite figures that arrive pre-attributed to their source - journalists academics use cases. Investors and quants (relevance 2) read country fundamentals as the backdrop against which company estimates get sanity-checked - investors quants use cases.
What should I know before requesting a sample?
Four honest caveats, stated plainly.
First, presence is not guaranteed. The graph imports a selection of what each provider publishes - not everything. Confirm a specific variable exists for your places before building a pipeline on it; that confirmation is precisely what a sample is for.
Second, scale is unnumbered. "Tens of thousands of variables across hundreds of provenance sources" is the honest ceiling of what the documentation states, and even that appears as coverage charts rather than published counts. Treat any precise total - ours included - as a floor.
Third, history is uneven. Depth rides with the facet: a line reaching to the 1960s in one country may begin a decade later in its neighbor. Longitudinal designs should read the declared bounds per series rather than assume a shared start.
Fourth, the grain is places, not firms. This is the macro layer of the vertical. When the question descends to named corporate groups, pair it with a filings or registry feed and join on geography and year - several sit on the same industry shelf.
Notes and related datasets
Notes that pair well with this page:
- Industrial conglomerates data hub - the pooled industry view, from filings and registries to this macro layer.
- World Bank WDI - Manufacturing, value added (% of GDP) - the single-indicator original, about 265 economies deep; the graph re-serves it beside tens of thousands of sibling variables.
- Eurostat Data Browser (Industrial Production & Structural Business Statistics) - the European production layer in official SDMX detail.
- BEA GDP by Industry - the United States production account, the domestic complement to the worldwide series.
- SEC EDGAR Company Facts & Submissions API - when the question drops from economies to filers; standardized XBRL for every US-listed conglomerate.
- Persona pages - what data-science, market-research, developer, journalist and investor teams each do with this slice.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
byVariable | map | Top-level result grouping keyed by statistical variable, so one pass returns every requested indicator side by side. | Manufacturing_ValueAdded_AsFractionOfGDP |
byEntity | map | Place-level grouping beneath each variable, keyed by entity identifier. | country/USA |
facetId | string | Numeric identifier of one provenance facet - the source attribution standing behind a measurement. | 18369491376878146239 |
date | string | Observation date, formatted per the variable's declared calendar. | 2020 |
value | number | The observed figure itself, typed on delivery. | 329484123 |
Additional fields on request | - | The provenance facet list, the observation arrays, per-facet observation counts and date bounds, and the pagination continuation handle - definitions and examples ship with your sample. | - |
Sample rows - shipped observations from the graph (one row per entity x variable x date x facet)
| Variable | Entity | Date | Value |
|---|---|---|---|
| Count_Person | country/USA | 2020 | 329484123 |
| Count_Person | country/DEU | 2020 | 83160871 |
| facet 18369491376878146239 | country/DEU | 2020 | 83160871 |
Coverage chips - geography, temporal range, granularity and scale
| Dimension | Coverage |
|---|---|
| Geography | Worldwide down to country, state and district/county level; United States coverage the most extensive |
| Temporal | Declared per variable and per provenance facet; multi-decade series common, census and World Bank lines back to the 1960s; each facet carries explicit earliest/latest bounds |
| Granularity | One observation per entity x statistical variable x date x provenance facet; no silent averaging across sources |
| Scale | Tens of thousands of statistical variables across hundreds of provenance sources; exact totals presented as coverage charts rather than published counts |
Questions buyers ask
How many statistical variables does Google Data Commons cover?
Tens of thousands across hundreds of provenance sources is the documented magnitude, and even that is presented as coverage charts rather than a published count - so treat any precise total, ours included, as a floor. The practical route is to name the indicators and places you need and let a sample confirm availability.
Whose statistics sit inside the graph?
Documented provenance includes the World Bank, OECD, Eurostat, the Food and Agriculture Organization, the World Health Organization, NOAA, the US Environmental Protection Agency, the Bureau of Labor Statistics, the Bureau of Economic Analysis and the U.S. Census Bureau - American Community Survey among others - alongside many national statistical offices. Each observation keeps its source attached as a facet rather than dissolving into an anonymous blend.
How far back does the data go?
It varies by variable and by source. Multi-decade series are common, and census and World Bank lines reach back to the 1960s. Every provenance facet declares its own earliest and latest dates, so available history is read per series instead of assumed from a site-wide claim.
Can the same indicator be compared across countries?
That is the design goal. One statistical variable resolves against many place entities in a single pass, so a manufacturing value-added share, a population count or an emissions figure arrives identically shaped for every market in scope - no per-country reformatting.
What exactly is one row of the data?
One observation: an entity (a place such as country/DEU), a statistical variable, a date and a provenance facet, carrying a typed value. Disagreement between sources survives as parallel facets rather than being averaged away, which keeps every downstream aggregate reproducible.
Does it contain firm-level data on conglomerates?
No. This is the macro layer of the industrial-conglomerates shelf: places and indicators, not companies. Pair it with a filings or registry feed keyed on geography and year when the analysis needs named corporate groups - several of those sit on the same shelf.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.