Datadory notebook

Cross-Industry Data Sources: 40% of the Catalog

1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.

1,744 datasets. Pick your catch.

How much of Datadory's catalog comes from multi-industry sources?

Datadory catalogs 1,744 datasets drawn from 1,124 unique sources across 163 cataloged industries, and only 149 of those sources appear in more than one industry. That 13.3% minority is the backbone of the catalog: multi-industry sources hold 703 of the 1,744 datasets, or 40.3%. Concentration rises steeply at the top. Twenty-nine sources span five or more industries and together account for 373 datasets, while the ten widest-reaching sources alone hold 216.

The remaining 1,041 datasets come from 975 single-industry sources, which makes the average multi-industry publisher (4.7 datasets) more than four times as productive as the average single-industry publisher (1.1). For buyers, that ratio reframes sourcing strategy: a handful of publishers decide whether dozens of verticals have usable data at all.

Which sources reach the most industries?

Eight publishers can be ranked on both industry reach and dataset volume, and the top of that list is entirely institutional statistics:

RankSourceIndustries servedDatasetsWhere it shows up
1U.S. Census Bureau2937Auto retail sales, building permits, county business patterns, five REIT sectors
2Eurostat2832Steel production indices, railway freight, EU hard-coal balances
3Hugging Face2427Advertising CTR benchmarks, 83 beer-review corpora, finance repositories
4Kaggle2225Ad-spend regression, gold prices, airline passenger-satisfaction survey
5data.gov (all name variants)60*76*Harvested federal, state and local open-data portals
6World Bank1719Tobacco prevalence, drinking-water access, Global Findex financial inclusion
7Yahoo Finance1315Quote pages covering mostly REIT verticals
8Nareit712T-Tracker quarterly operating performance across REIT sectors

*data.gov appears in the catalog under uppercase and lowercase variants; counted together its entries reach 60 industries with 76 datasets, and Census-named sources reach 47 industries with 65 datasets.

Government depth extends past the podium: Statistics Canada serves 10 industries, the SEC 9 and UK ONS 6, though the catalog analysis reports their industry reach without dataset totals. Brief-level citations confirm the hierarchy: across the 159 industry briefs Datadory publishes, Eurostat is named in 56, Data.gov in 51, Kaggle in 49, Hugging Face in 47, the World Bank in 45 and the Census Bureau in 34.

What kinds of publisher span multiple industries?

Four families dominate the 149-source tier.

Statistical agencies lead. The U.S. Census Bureau tops the table at 29 industries and 37 datasets. Monthly Retail Trade Survey rows feed automotive-retail datasets, the Building Permits Survey anchors homebuilding, County Business Patterns and the Economic Census quantify health-care-supplies manufacturing, and five of its industries are REIT sectors, including hotel-resort-reits via the economic indicators briefing room. Eurostat reaches 28 industries and 32 datasets, from steel production indices to railway freight and hard-coal balances.

Machine-learning hubs come next. Hugging Face spans 24 industries and 27 datasets: advertising CTR benchmarks, 83 beer-review corpora for brewers, and roughly 1,352 finance repositories for specialized finance. Kaggle spans 22 industries and 25 datasets, from ad-spend regression to gold prices and the airline passenger-satisfaction survey.

Financial platforms form the third family. Yahoo Finance quote pages cover 13 mostly REIT industries with 15 datasets, and Nasdaq reaches 6.

Sector bodies close the gap. Nareit serves 7 REIT verticals with 12 datasets, and the EIA bridges 7 energy-linked industries, landing CBECS building-energy tables in office-reits.

Why do thin industries run on shared infrastructure?

Because few publishers bother with narrow verticals directly. Health-care-supplies draws 92 of its 93 datasets from multi-industry sources; diversified-support-services draws 109 of 115 (95%), construction-engineering 115 of 122 (94%), and hotel-resort-reits 106 of 113 (94%). Remove the shared agencies and those categories effectively go dark.

The pattern cuts both ways. A buyer screening health-care supplies, construction engineering or diversified support services is really evaluating Census, Eurostat, World Bank and Kaggle feeds wearing different labels. Vendors selling "industry-specific" packages for these verticals are usually repackaging the same public series, so differentiation lives in cleaning, geography joins and delivery, rarely in exclusive collection.

What does one methodology buy you across borders?

That consistency matters for anything feeding a model or a market-sizing deck: definitions, units and revision cycles stay stable when the series crosses industry boundaries. Stitching together 20 sector-specific vendors produces the opposite, a patchwork of vintages and definitions that needs constant reconciliation.

The same logic holds inside one publisher's family: tobacco prevalence, drinking-water access and the Global Findex's financial-inclusion measures all arrive through World Bank indicator APIs under shared terms.

Geography follows the same publisher logic. Catalog-wide, 597 of 1,744 datasets (34.2%) center on United States coverage and 586 carry global or multi-country scope, with Europe/EU accounting for another 125. Those buckets map onto the shared tier almost one-to-one: national agencies such as the Census Bureau, Statistics Canada and UK ONS produce the country-deep series, while the World Bank produces the country-wide ones. A buyer assembling international coverage is therefore choosing between two publisher families, not sampling hundreds of vendors.

What should buyers watch out for with shared sources?

Three cautions temper the story.

Publisher names fragment. The Census Bureau appears under 14 name variants and data.gov under 34, so coverage assessments made on raw source names undercount badly. Normalize names, or deduplicate by publisher identity, before drawing conclusions.

Some overlap is a tagging artifact. The self-storage-reits brief explicitly excludes seven cross-tagged semiconductor datasets, and data-processing-outsourced-services carries five viticulture datasets that serve wine research rather than outsourcing buyers. Cross-tags indicate relevance, not fit.

What this means for you

Five moves follow directly from the numbers.

  1. Strategy and research analysts: size adjacent markets from publisher coverage rather than SIC codes. If health-care-supplies depends on shared agencies for 92 of 93 datasets, the addressable data surface equals those agencies' joint geography and periodicity.
  1. Procurement teams: audit overlap before renewing sector vendors. When 149 publishers explain 703 of 1,744 datasets, duplicate subscriptions to repackaged public series are common and provable line by line.
  1. Journalists and academics: attribute thin-vertical statistics to the underlying agency, since 94% or more of datasets in categories like hotel-reits or construction-engineering originate upstream of any industry brand.

Entry points matched to each role: investors and quants, market researchers, journalists and academics, data scientists, and sales growth teams mapping territories across verticals.

The widest-reaching sources in Datadory's 1,744-dataset catalog
RankSourceIndustries servedDatasetsRepresentative coverage
1U.S. Census Bureau2937Retail trade, building permits, business patterns, five REIT sectors
2Eurostat2832Steel production indices, rail freight, hard-coal balances
3Hugging Face2427Advertising CTR benchmarks, beer-review corpora, finance repos
4Kaggle2225Ad-spend regression, gold prices, airline satisfaction survey
5data.gov (all variants)60*76*Federal, state and local portal harvest
6World Bank1719Tobacco prevalence, water access, Findex inclusion
7Yahoo Finance1315REIT-heavy quote pages
8Nareit712T-Tracker REIT operating performance

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Which data sources cover the most industries?

The U.S. Census Bureau leads with 37 datasets across 29 industries, ahead of Eurostat at 32 datasets across 28 industries. Hugging Face spans 24 industries and Kaggle 22. Counted across name variants, data.gov-named sources contribute 76 datasets spanning 60 industries, while the World Bank reaches 17 industries, Statistics Canada 10, the SEC 9 and UK ONS 6.

How much of a niche industry's datasets come from shared sources?

Almost all of it in thin verticals. Health-care-supplies draws 92 of its 93 datasets from multi-industry sources, construction-engineering 115 of 122, diversified-support-services 109 of 115 and hotel-resort-reits 106 of 113. Only one health-care-supplies dataset comes from a publisher exclusive to that industry.