Catalog statistics

How Datadory Validates Every Dataset in Its Catalog

As of August 2026 · computed on 2026-08-25

datasets
1,744
industries
144
sources
1,055
top coverage
66

Catalog validation at a glance

Every number on this page comes from the same audit pass over Datadory's master catalog, computed programmatically on 2026-08-22 and stated here as of August 2026. The unit of accounting is always the full denominator: 1,744 datasets drawn from 1,124 unique sources across 163 cataloged industries. Nothing on this page is sampled or extrapolated — each row below is a direct count over all 1,744 records, so any figure can be reproduced from the catalog itself.

How strict is the 10-point quality rubric?

The rubric runs on a 10-point scale, and the catalog-wide average lands at 7.81 out of 10. Scores skew high: 1,096 of 1,744 datasets (62.8%) score 8 or higher, and 145 records (8.3%) sit at the ceiling of 10. A low score does not remove a dataset from the catalog — 30 records (1.7%) score 4 or below, mostly in thin industries, and they are listed with their score attached so buyers can price the risk themselves. Across 24 industry briefs, government-named top datasets average 8.81 versus 8.47 for the rest, and 16 of the 22 ceiling-scored top datasets named in those briefs are government or public-science sources such as openFDA, NCBI GenBank, PubChem and ClinicalTrials.gov.

Where does the validated catalog come from geographically?

Geographic scope is bucketed from each record's coverage text using keyword normalization, and the result is nearly split down the middle: 597 of 1,744 datasets (34.2%) center on United States coverage versus 586 (33.6%) with global or multi-country scope, followed by Europe/EU at 125 (7.2%), the United Kingdom at 74 (4.2%), India at 32, Canada at 27, China at 11 and Australia at 8. Provenance skews institutional: 781 datasets (44.8%) are public_dataset type, and the catalog's 1,744 records compress into just 1,124 unique sources. Reach concentrates further — 703 datasets (40.3%) come from 149 sources serving more than one industry, led by the U.S. Census Bureau at 37 datasets across 29 industries, Eurostat at 32 across 28, Hugging Face at 27 across 24 and Kaggle at 25 across 22.

Methodology

All figures on this page were computed programmatically from the 1,744-record master catalog on 2026-08-22; the page reads as of August 2026. Scope is the entire catalog — 1,744 datasets, 163 cataloged industries, 1,124 unique sources — with no sampling and no exclusion of failed or low-scoring records. Percentages are taken against the full 1,744 denominator unless a narrower base is stated, and denominators are never rounded away. Geographic buckets apply keyword normalization to the free-text coverage.geographic field, so mixed-scope records land in other/mixed rather than being forced into a single country. Format counts draw on a raw by_format field containing 1,061 distinct labels — json (api), json-stat, json-ld and sdmx-json variants among them — which pipelines must normalize before comparing; the CSV (734) and JSON (675) counts above follow that normalization. Quality uses the integer-valued 10-point rubric, with observed scores running 4 through 10 and the mean reported to two decimals (7.81).

Datadory catalog validation at a glance (as of August 2026, computed 2026-08-22)
MetricValueDenominatorDetail
Datasets cataloged1,744Full catalogFrom 1,124 unique sources
Field definitions verified1,495 (85.7%)1,744 datasets249 records lack confirmed schema docs
Average quality score7.81 / 101,744 datasetsScores observed 4–10
Records scoring 8+1,096 (62.8%)1,744 datasets145 score a perfect 10
Viable industries159163 catalogedViability bar: 3+ datasets
Quality score distribution across the 1,744 cataloged datasets (10-point rubric)
Quality scoreDatasetsShare of catalog
101458.3%
953430.6%
841723.9%
729617.0%
623013.2%
5925.3%
4 or below301.7%
Geographic buckets after keyword normalization of coverage text
Coverage bucketDatasetsShare of catalog
United States59734.2%
Global / multi-country58633.6%
Other / mixed28416.3%
Europe / EU1257.2%
United Kingdom744.2%
India321.8%
Canada271.5%
China110.6%
Australia80.5%

Questions worth asking

Does Datadory remove datasets that score badly?

No. A quality score flags risk rather than triggering removal: 30 of the 1,744 cataloged datasets (1.7%) score 4 or below on the 10-point rubric and remain listed with their scores attached, concentrated in thin industries. The catalog-wide average is 7.81 out of 10, and 1,096 records (62.8%) score 8 or higher.

How many industries does the validated catalog cover?

Datadory catalogs 163 industries, of which 159 clear the minimum viability bar of 3+ datasets; 11 viable industries still hold fewer than 8. Four cataloged industries have no primary or secondary coverage at all: real-estate-development, real-estate-services, data-center-reits and other-specialized-reits.

How current are these validation figures?

Every figure was computed programmatically from the 1,744-record master catalog on 2026-08-22, so the page reads as of August 2026. Counts cover the full catalog with no sampling, and geographic buckets reflect keyword normalization of each record's free-text coverage field.

Cite this

Every figure here traces to one programmatic rollup of the Datadory catalog, so quote it with the snapshot attached: “of the 1,744 datasets Datadory catalogs as of August 2026…”. Keep the denominator next to the percentage — 1,744 of an unspecified universe is a different claim.

API, files, or your warehouse. Daily, weekly, or hourly.

Read the rest of the rollup