Datadory notebook

How Data Scientists Use Other Specialty Retail Data

Datadory delivers other specialty retail data covering monthly NACE G47 turnover indices across 41 European geographies back to January 1991, roughly 171 million agri-trade rows over 245+ countries since 1961, bilateral HS-chapter flows worth $63.8 billion in chapter 16 alone, ~356,000 packaged-food product records with nutrition grades, recall event history and company benchmarks - delivered daily, weekly, or hourly - your call.

1,744 datasets. Pick your catch.

What does a data scientist actually get from other specialty retail data?

Open any modeling notebook in this vertical and the constraint is not ideas - it is labeled history at a usable grain. This slice solves it unevenly but well: two sources carry deep time series (Eurostat back to January 1991, FAOSTAT back to 1961), one ships an ML-ready product table (~356,000 rows), and the rest are feature layers you join on geography, HS code or barcode.

Nine of the twenty pooled records clear Datadory's relevance bar for data scientists, scored 9 down to 6 with a mean of 7.2 against a catalog-wide average of 7.81. The gap reflects persona fit rather than raw coverage - the nutrition databases lose ground on how narrowly their records map to modeling workflows, not on what they document.

Which dataset anchors a European demand panel?

Eurostat - European Statistical Office Data Portal is the anchor, and its sts_trtu_m series ('Turnover and volume of sales in wholesale and retail trade') is dimensioned by business indicator (NETTUR net turnover, VOL_SLS volume of sales), 38 NACE Rev.2 codes spanning the full G47 retail tree down to G477 clothing and G473 automotive fuel, seasonal adjustment (NSA, CA, SCA), unit (I21 index 2021=100, PCH_PRE and PCH_SM percentage changes) and 41 geographic codes from EU27_2020 aggregates to EFTA and candidate countries.

Each monthly vintage exposes 1,483 dimension-combination rows - 38 NACE codes x 5 units x 3 adjustments x 41 geographies across 424 months - with June 2026 as the newest observation (EU27 G47 net turnover index 122.2, NSA). That is a demand panel reaching back to January 1991 at monthly grain, deep enough to fit seasonality, turning points and country-level shocks without splicing sources.

Which table do you train product-level models on?

Two caveats belong in your data card. First, the mirror is a static version-5 snapshot published 2017-09-18 - the parent project now cites over four million products and republishes continuously, so treat the Kaggle file as a frozen benchmark for reproducible baselines rather than current state. Second, completeness is uneven: 67,000 fully populated products out of ~356,000 means missingness-aware column selection decides more of your feature set than taste does.

What event and discovery layers feed the feature store?

Three government layers round out the panel. FDA Recalls, Market Withdrawals & Safety Alerts - Searchable Listings carries eight columns per event (Date, Brand Name(s), Product Description, Product Type, Recall Reason Description, Company Name, Terminated Recall indicator, Excerpt) with roughly 1,025 notices live in a rolling three-year window - a ready-made label source for brand-level risk scoring. Deeper history sits in the openFDA food-enforcement export: 29,310 records reaching back to 2004-era event IDs, enough span to estimate hazard rates rather than eyeball a notice list.

Data.gov Retail Catalog (552k+ Datasets Federated Search) is your discovery index: a federated walk of its 552,271-record catalog surfaced ~264 'retail' datasets, including Washington State taxable retail sales by county and NAICS back to 1994 and CDC's Modified Retail Food Environment Index. The keyword='food' slice adds 208 more records (30 tagged 'meat'), from FSIS inspection directories to city grocery files - jurisdiction-hopping context you cannot get from any single national statistical office.

Can you benchmark company performance without a data vendor?

As a modeling input it comes with limits. Ranking tables arrive as browsable HTML rather than a flat export, archives reach back to only about 2020, and Kantar explicitly notes its compilation is built from public disclosures and may not match company filings. Treat it as a benchmarking cross-check and calendar-alignment reference (the 4-5-4 PDFs solve fiscal-week comparability). The July 2026 Retail Monitor reading - core sales up 0.30% month over month and 4.72% year over year - lands 10-14 days after month end without revisions, which makes it a useful early signal next to Eurostat's official indices.

Which datasets make the ranked modeling shortlist?

Ranked by Datadory's persona relevance for data scientists (2 = directly serves modeling workflows), then by quality score:

How do you assemble these sources into one pipeline?

A step-by-step sketch that reaches one joined analytical table:

Where should you go next?

The other-specialty-retail data guide is the pillar: all 20 pooled records for the industry - 10 primaries plus 10 cross-tagged neighbors - scored and grouped by workflow. For the page-level version of this shortlist, best data for other specialty retail teams ranks every qualifying dataset with coverage and field notes for your role.

Deeper dives cover the two anchors: european retail sales data by country expands on the G47 index family for cross-country forecasting work, and the best other-specialty-retail datasets list plus free other-specialty-retail datasets roundup turn the shortlist into a working queue.

Ranked other-specialty-retail datasets for data scientists (by Datadory persona relevance, August 2026)
RankDatasetModeling useHistoryCoverage
1Eurostat - European Statistical Office Data PortalMonthly NACE G47 turnover/volume indices; SCA-adjusted exogenous regressorsJanuary 1991 through June 2026, 424 monthly vintagesEU27 aggregates, member states, EFTA and candidates; 1,483 dimension combinations per vintage
4Kaggle - World Food Facts (Open Food Facts mirror)~356,000 products x 150+ columns; Nutri-Score and PNNS targets for MLStatic version-5 snapshot published 2017-09-18Ingredients, allergens, additives, per-100g nutrition, packaging, brands
5Data.gov Retail Catalog (552k+ Datasets Federated Search)Discovery index across a 552,271-record federal catalogWashington State taxable retail sales by county and NAICS back to 1994~264 'retail' records incl. CDC Modified Retail Food Environment Index
6Data.gov - Food-Tagged Open Datasets CollectionFood-keyworded discovery layer for jurisdiction-hopping contextVaries per harvested record208 records, 30 tagged 'meat', incl. FSIS inspection directories and city grocery files
7FDA Recalls, Market Withdrawals & Safety Alerts - Searchable ListingsRecall-event labels for brand-level risk scoringRolling three-year window live; 2004-era event IDs in the enforcement exportEight columns: date, brands, product description, type, reason, company, termination, excerpt
8National Retail Federation (NRF) Research & Insights HubCompany benchmarks (Top 100 + Hot 25 = 125 firms) and 4-5-4 calendar alignmentRankings to about 2020; calendars to the 2009-2011 cycleJuly 2026 Retail Monitor core sales +0.30% m/m, landing 10-14 days after month end
9Nutritionix Natural Language & Grocery APILive UPC enrichment for basket-level and price-per-nutrient featuresCurrent-state item snapshots rather than archival history1,053,256 barcoded grocery items across 48,317 brands; stated match rate above 92% for US/Canadian scans

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Other Specialty Retail Europe - 41 geographies: EU27_2020 and euro-area aggregates…

Eurostat - European Statistical Office Data Portal Data

Other Specialty Retail 245+ countries and territories across all FAO regional groupings

FAOSTAT Food and Agriculture Statistics

further reshapes on request

Other Specialty Retail Worldwide bilateral flows between all reporting economies, at…

Meat & Seafood Preparations Trade Profiles Data (HS16)

Other Specialty Retail Global products from contributing countries, weighted toward…

Kaggle - World Food Facts (Open Food Facts Mirror)

Other Specialty Retail United States: national agencies plus harvested state, county…

Data.gov Retail Catalog (552k+ Datasets Federated Search)

Other Specialty Retail United States - federal agencies plus participating state…

Data.gov - Food-Tagged structured datasets Collection

further fields on request …+9 more

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Which dataset anchors a European retail demand panel?

Eurostat's sts_trtu_m series: monthly NACE G47 net turnover (NETTUR) and volume-of-sales (VOL_SLS) indices for EU27 aggregates, every member state, EFTA members and candidate countries - 41 geographies - from January 1991 onward, cut by 38 NACE Rev.2 codes down to G477 clothing and G473 automotive fuel. Each vintage holds 1,483 dimension combinations across 424 months, and June 2026 stands as the newest observation at an EU27 G47 net-turnover index of 122.2.

What can you model with packaged-food product data?

The Kaggle-hosted World Food Facts mirror of Open Food Facts: ~356,000 products (67,000 complete per the version note) across 150+ columns covering ingredients_text, allergens, additives_n, per-100g nutrition, packaging, brands, nutrition_grade_fr Nutri-Score and pnns_groups_1/2 - with 551 public notebooks evidencing how often it seeds baselines. It is a static September 2017 snapshot while the parent project has passed four million products, so treat it as a frozen benchmark rather than current state.

Which sources carry recall history worth joining?

Two layers of the same FDA record set. The searchable listings view holds roughly 1,025 notices in a rolling three-year window across eight columns (date, brand names, product description, type, recall reason, company, terminated-recall flag, excerpt), while the openFDA food-enforcement export extends to 29,310 records reaching back to 2004-era event IDs. Together they give brand-level risk models both a live signal and a decade-plus training base.