Datadory notebook

Nutrition Database for App Developers: Comparing the Four Sources That Ship

Datadory delivers other specialty retail data covering item-level nutrition records, barcode coverage, allergen and diet labels and federal recall notices - delivered daily, weekly, or hourly. Four catalog records anchor an app build: Nutritionix Natural Language & Grocery API at 1,266,570 items including 1,053,256 barcoded grocery items across 48,317 brands, the Edamam Food Database API near 1,000,000 items with roughly 790,000 UPCs, the Kaggle World Food Facts mirror at about 356,000 packaged foods across 150+ columns, and FDA recall listings holding roughly 1,025 live notices for safety screening.

1,744 datasets. Pick your catch.

Which nutrition databases can an app actually ship on?

Of the ten primary datasets Datadory catalogs for Other Specialty Retail, exactly three carry item-level nutrition data a consumer app can query. The other seven primary sources measure retail trade or trade flows - Eurostat's NACE G47 indices, FAOSTAT's ~171 million rows of production and trade series, the NRF's company rankings - none of which answer "what is in this barcode?"

The three that matter to builders:

  1. Nutritionix Natural Language & Grocery API - 1,266,570 items, including 1,053,256 barcoded grocery items across 48,317 brands and 202,837 restaurant items from 860 chains.
  2. Edamam Food Database API - close to 1,000,000 food items with roughly 790,000 unique UPC/barcodes plus natural-language parsing.
  3. Kaggle - World Food Facts (Open Food Facts mirror) - ~356,000 packaged foods across 150+ columns, frozen at its September 2017 snapshot.

A fourth input belongs next to any of them rather than inside them: the FDA Recalls, Market Withdrawals & Safety Alerts listings hold roughly 1,025 live notices in a rolling three-year window, backed by the openFDA food enforcement export's 29,310 historical records. A nutrition feature that surfaces allergens without checking recalls is half a feature.

What fields does each nutrition source return?

Nutritionix Natural Language & Grocery API returns per-item records with per-serving nutrient values drawn from a dietitian-verified base; coverage is deepest for the United States and Canada, with international common-foods and multi-language views layered alongside.

Edamam Food Database API splits its close-to-one-million items into roughly 790,000 UPC/barcodes, about 130,000 branded restaurant items and around 100,000 common foods, and generates diet, allergy and nutrition labels over the top - the same label layer that powers diet-label filters in third-party apps. Records support per-serving or per-100g nutrient bases, so you can normalize against package weights client-side.

The Kaggle World Food Facts mirror is the only one of the three that hands you raw columns instead of a response object: product names, brands, categories, ingredients, allergens and per-100g nutrition across 150+ columns in CSV, TSV and SQLite. Coverage skews toward France and other European markets at the 2017 snapshot, and only 67,000 of the ~356,000 products are flagged complete.

Which nutrition dataset holds up as an ML training corpus?

One does, and it is why ML teams reach for the mirror over both lookup services. The Kaggle World Food Facts extract is 1,010,256,825 bytes across CSV and SQLite files, published as version 5 on 2017-09-18 and unchanged since. Static is a defect for a lookup service and a virtue for a training corpus: rows do not move between runs, and every column - product names, brands, categories, ingredients, allergens, per-100g nutrients - arrives labeled.

For demand- and market-side features that surround a nutrition tracker, two complements round out the stack. FAOSTAT Food and Agriculture Statistics ships 69 bulk domains totalling ~171 million rows covering 245+ countries since 1961 - food balance sheets and production volumes contextualize what your users log across six decades.

How do you add recall screening to a nutrition feature?

A three-step pattern keeps the safety layer separate from the nutrition layer:

  1. Resolve the scan against Nutritionix or Edamam to get brand, product description and UPC.
  2. Screen the identifiers against recall data: the openFDA food enforcement export holds 29,310 historical records, while the FDA's searchable listings hold roughly 1,025 live notices in a rolling three-year window.
  3. Match on brand and product text, because recall notices describe products in regulatory language rather than barcodes, then surface the notice date and classification alongside the nutrition panel.

The same pattern extends beyond the US border only partially: neither Nutritionix nor Edamam carries recall status, and no European equivalent is cataloged in this slice, so EU launches need their own safety feed.

Where to go next

The other specialty retail data guide is the pillar for this industry: all twenty pooled datasets scored and grouped by workflow, with the industry hub mapping the same ten primary and ten secondary records. To see how these sources rank against the rest of the slice, the best other-specialty-retail datasets ordering covers the catalog, and the developers-builders persona page collects the same four sources into a build workflow - the persona page exists because developers-builders is one of five personas Datadory maps against this industry.

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Other Specialty Retail United States and Canada primary depth (>92% stated UPC match)

Nutritionix Natural Language & Grocery API

16 documented spine fields · with fuller nutrient vectors …+13 more

Other Specialty Retail Global product coverage with US-centric depth on grocery…

Edamam Food Database API

Other Specialty Retail Global products from contributing countries, weighted toward…

Kaggle - World Food Facts (Open Food Facts Mirror)

Other Specialty Retail 245+ countries and territories across all FAO regional groupings

FAOSTAT Food and Agriculture Statistics

further reshapes on request

Other Specialty Retail United States, with affected states noted in per-record…

FDA Recalls, Market Withdrawals & Safety Alerts - Searchable Listings

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

What is the best nutrition database for app developers?

For live lookups, Nutritionix Natural Language & Grocery API is the largest at 1,266,570 items with 1,053,256 barcoded grocery items across 48,317 brands, while the Edamam Food Database API covers close to 1,000,000 items including roughly 790,000 UPCs. For a static corpus, the Kaggle World Food Facts mirror offers about 356,000 products across 150+ columns.

Which nutrition database includes allergen and diet labels?

The Edamam Food Database API generates diet, allergy and nutrition labels automatically across its near-1,000,000-item database, covering attributes such as allergen and diet compliance alongside nutrients returned per serving or per 100g. Nutritionix provides dietitian-verified records with per-serving nutrient values, and the Kaggle mirror exposes ingredient and allergen columns you must parse yourself.

Which nutrition dataset is stable enough for ML training?

The Kaggle World Food Facts mirror: published as version 5 on 2017-09-18 and unchanged since, about 356,000 products across 150+ columns with 67,000 flagged complete. Rows do not move between runs, which makes it dependable as a benchmark corpus even though a lookup service needs fresher coverage.

Can a nutrition feature screen for recalls too?

Yes. The FDA Recalls, Market Withdrawals & Safety Alerts listings hold roughly 1,025 live notices in a rolling three-year window, backed by the openFDA food enforcement export's 29,310 historical records. Match on brand and product text, because recall notices describe products in regulatory language rather than barcodes, then surface the notice date and classification alongside the nutrition panel.