Glossary

Crowdsourced Food Database

A crowdsourced food database is assembled from volunteer product scans rather than official surveys, so coverage concentrates where scanning communities are active. Open Food Facts holds about 4.7 million packaged products with ingredients, Nutri-Score, NOVA and packaging fields, updated daily under commercial delivery terms 1.0.

What is Crowdsourced Food Database?

The unit of contribution is a barcode scan plus a photo of the label; volunteers type ingredients, brands and categories, and algorithms derive scores such as Nutri-Score and NOVA processing class. Open Food Facts accumulates roughly 4.7 million barcoded products since its 2012 launch, one row per barcode/product, with created and last-modified timestamps back to launch.

Coverage is world-shaped but unevenly so: strongest in France, the US, Spain, Germany and Italy where scanning communities are most active. Delivery matches the crowd model - nightly bulk exports (CSV ~0.9 GB compressed / ~9 GB uncompressed across 211 columns, plus MongoDB dump and JSONL) and a live JSON API whose usage policy maps one call to one real user scan. Licensing splits three ways: commercial delivery terms 1.0 for the database, Database Contents License for individual contents, commercial delivery terms-SA 3.0 for product images.

Why does Crowdsourced Food Database matter when choosing a dataset?

Crowd-built corpora trade completeness control for scale at zero cost. Buyers who treat them as authoritative registries discover the trade-off in production.

  • Coverage mirrors volunteer geography. A European grocery analysis works well; a Southeast Asian shelf audit will find thin rows.
  • Field completion is voluntary. Of 211 columns, many are sparse per product - profile fill rates before joining nutrition claims onto them.
  • Share-alike follows the database. commercial delivery terms obligations attach to derived databases, so a commercial app blending these rows inherits attribution and share-alike duties.
  • API etiquette is contractual. The one-call-per-real-scan policy exists because complete nightly exports already serve bulk needs - scraping the API is both discouraged and unnecessary.
  • Freshness is real but per-record. Daily updates apply to changed records, not to every stale row.

How do you evaluate Crowdsourced Food Database in a data source?

  1. Measure field density for your category before committing. The 211-column CSV export spans nutrients, labels and packaging; run a fill-rate query on your categories instead of trusting headline product counts.
  1. Check geographic skew against your market. Open Food Facts states coverage is strongest in France, the US, Spain, Germany and Italy - other markets ride on sparser scanning activity.
  1. Confirm which license layer governs your use. commercial delivery terms 1.0 (database), Database Contents License (contents) and commercial delivery terms-SA 3.0 (images) apply simultaneously; attribute each correctly.
  1. Match access mode to volume. Point lookups suit the live API; bulk work belongs on the ~0.9 GB compressed CSV or MongoDB dump - see nightly-full-export.
  1. Plan around timestamp-based deltas. Records carry created/last-modified timestamps back to 2012, so incremental sync can key off modified dates rather than full reloads.

Where to see it in context: nightly-full-export.

Entries that sit next to Crowdsourced Food Database in this glossary:

Frequently asked questions

How big is the Open Food Facts database?

About 4.7 million packaged food products worldwide, distributed as nightly exports - CSV of roughly 0.9 GB compressed / 9 GB uncompressed across 211 columns - plus a live read API.

Why does crowdsourced food data have country gaps?

Because records originate from volunteer scans; Open Food Facts itself notes coverage is strongest in France, the US, Spain, Germany, Italy and other countries with active scanning communities.

Is Open Food Facts free for commercial apps?

The data is openly licensed but not obligation-free: commercial delivery terms 1.0 requires attribution and share-alike for derived databases, contents fall under the DbCL, and API calls should correspond to real user scans.

Datasets containing this field

Datasets containing Crowdsourced Food Database

6 datasets carry crowdsourced food database in the catalog. Open one, count the fields, judge for yourself.

Packaged Foods & Meats

Open Food Facts - Open Database & API

Packaged Foods & Meats United States, with state-level… · 2004 to the present

openFDA Food Enforcement (Recalls) API Data

Packaged Foods & Meats United States-focused grocery products · Current catalog snapshot

Spoonacular Food API - Product Search

badges · importantBadges · ingredients …+2 more

Packaged Foods & Meats

USDA AMS Market News - Livestock, Poultry & Grain

Packaged Foods & Meats

USDA ERS Food Availability (Per Capita) Data System

Packaged Foods & Meats United States national plus all 50… · Current annual series from 1997

USDA ERS Food Expenditure Series

Every listing shows the field dictionary, sample rows, and coverage before you commit. API, files, or your warehouse. Daily, weekly, or hourly.

Get sample rows