Other Specialty Retail · Nutritionix

Nutritionix Natural Language & Grocery API

Datadory delivers other specialty retail data covering the verified-nutrition layer of the grocery shelf: the Nutritionix Natural Language & Grocery API packages 1,266,570 food items - 1,053,256 barcoded grocery products across 48,317 brands, 202,837 restaurant items from 860 chains and a dietitian-built common-foods layer - as typed rows keyed on item identifier and UPC, with per-serving nutrient panels attached. Delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

What is the Nutritionix Natural Language & Grocery API?

It is the verified-nutrition layer of the American and Canadian food supply, published as queryable records. Nutritionix - now owned by Syndigo - operates what it describes as the largest verified nutrition database: 1,266,570 food items, of which 1,053,256 are barcoded grocery items across 48,317 grocery brands, with a stated match rate above 92% for US and Canadian UPC scans. Around them sit 202,837 restaurant items from 860 chains monitored for menu changes, and a curated common-foods layer where an in-house registered-dietitian team maps 10,477 tags onto 38,828 everyday food phrases built on top of USDA data.

Three query families cover the corpus. Natural-language parsing converts a sentence - 'two scrambled eggs and toast' - into structured per-food nutrient rows. Instant search returns autocomplete candidates as you type. Direct lookup resolves a single item by UPC barcode or internal identifier. Scale tells you how battle-tested the plumbing is: roughly 20,000 health apps push about 250 million queries through it every month. On the Other Specialty Retail shelf this record owns the item-level nutrition lane, and it scores 6/10 on our rubric - carried by breadth and verification, docked because response-shape documentation is captured from observed structure rather than a fully published schema.

What do sample rows look like?

Three shapes cover the three query families:

# natural-language parse - one sentence becomes per-food rows
query_text     : "2 scrambled eggs and 1 slice wheat toast"
food_name      : Scrambled Eggs, 2 large
serving_qty    : 1          serving_unit: serving
nf_calories    : 199        nf_total_fat: 15.0 g
nf_protein     : 13.0 g     nf_sodium: 336 mg

# grocery record keyed by UPC - the enrichment join key
upc            : 041631000564
food_name      : Wheat Bread
brand_name     : Nature's Own
serving_qty    : 1          serving_unit: slice
serving_weight_grams : 43
nf_calories    : 110        nf_total_carbohydrate: 20.0 g
nf_dietary_fiber: 3.0 g     nf_sugars: 2.0 g

# restaurant row from a monitored chain menu
food_name      : Grilled Chicken Sandwich
brand_name     : Chick-fil-A
serving_qty    : 1          serving_unit: sandwich
nf_calories    : 320        nf_sodium: 970 mg

Read the anatomy rather than the digits. The parse block shows the grain decision the platform makes for you: one sentence, two foods, two separate rows, each carrying its own serving definition. The grocery block is the workhorse for retail enrichment - UPC as the join key, a named brand, and a per-serving panel anchored by serving_weight_grams so household measures convert cleanly. The restaurant block shows the second shelf living in the identical schema, which is why one pipeline serves both. Your sample ships cut against live records in these shapes before anything is committed.

Which fields does the field dictionary define?

Sixteen fields form the item spine, grouped four ways. Identity: food_name and brand_name, the latter splitting the 48,317 grocery brands from the 860 monitored chains. Serving block: serving_qty, serving_unit and serving_weight_grams - the trio that makes panels comparable, because 110 calories per slice and per loaf are different claims. Nutrient panel: the eight nf_ values from nf_calories through nf_potassium, reported per stated serving in grams and milligrams. Identifiers: nix_item_id, the internal key that survives renames and reformulations and rides alongside UPC for exact-item resolution.

Everything beyond the spine arrives through the request note rather than being invented into the schema: the full nutrient vector behind each nf_ summary value with its derivation attributes, the brand taxonomy with grocery categorization and derived wellness claims, multilingual item names across 15+ languages, and the common-foods phrase table mapping 38,828 expressions to 10,477 dietitian tags. Name the extras your workflow touches when you request a sample and they land as populated columns, not promises.

How wide does coverage run, and at what grain?

Geography - United States and Canada carry the depth that matters for barcoded goods: the stated above-92% UPC scan match rate applies to North American shelves. International extension comes through common foods and the multi-language add-on, not through foreign grocery UPC depth, so a European private-label barcode should be tested in your sample rather than assumed.

Temporal - the corpus is continuously maintained while items stay current. More than 3,000 grocery items are added or updated monthly, and the 860 restaurant chains are watched for menu changes, so reformulations and new SKUs arrive as updates rather than waiting for an annual re-release. There is no deep historical archive here; this is a current-state shelf, not a longitudinal study series.

Granularity - strictly item-level, one record per food with its per-serving nutrient panel, resolvable three ways: by UPC barcode, by internal item identifier, or by free-text phrase through the parser and autocomplete. Against the wider Datadory catalog - 1,744 datasets averaging 7.81 - this record scores 6/10, with breadth and verification pulling up and inferred response documentation holding it down.

How is the data delivered through Datadory?

API, files, or your warehouse. Daily, weekly, or hourly.

Channel and cadence are settings, not projects. Files suit the team loading the whole grocery corpus once and joining it to their own product tables offline. A direct pipe suits warehouses running assortment screens in SQL next to sales data. Record-level feeds suit anything interactive - a scanner page, a logging assistant, a chatbot answering 'how much sodium is in this?' mid-conversation. Every delivery carries the field dictionary above unchanged plus validation rows, units stay native to the source (calories, grams and milligrams per stated serving) unless you ask for conversions, and column naming locks at sample time.

Who builds on this data, and for what?

  • Product enrichment at barcode speed - e-commerce and retail-media teams attach complete nutrient panels to listings by UPC lookup instead of transcribing labels by hand.
  • Conversational food logging - health apps turn plain sentences into structured nutrient rows via natural-language parsing, the mechanic behind roughly 250 million queries a month.
  • Assortment and claim screening - category managers filter the grocery shelf on sugar, sodium, fiber and protein thresholds per serving to assemble better-for-you sets.
  • Menu-change intelligence - analysts track nutrition drift across 860 monitored restaurant chains without re-collecting a single menu board.
  • Model training - ML teams fit parsers and classifiers on a million-plus labeled, consistently keyed food records.
  • Reference citation - researchers and journalists anchor claims about branded and restaurant nutrition on a dietitian-verified corpus rather than crowd-sourced guesses.

Which datasets pair with this one?

  • Edamam Food Database API - the nearest competitor: close to a million items, about 790,000 of them UPC-barcoded. This record wins on verified depth and restaurant coverage; Edamam wins on diet-and-recipe add-ons.
  • World Food Facts (Open Food Facts mirror on Kaggle) - roughly 356,000 volunteer-contributed packaged foods with 150+ columns; crowd-sourced breadth beside verified panels, useful as a cross-check layer.
  • Food-Tagged Collections on Data.gov - the government-statistics complement: 208 food-keyed federal catalog records for demand and policy context rather than item-level labels.
  • NRF Research & Insights Hub - industry-level retail benchmarks to sit above the item-level shelf this record provides.
  • The Other Specialty Retail hub holds all ten records in the slice, including USDA FoodData Central's branded-foods files on the adjacent packaged-foods shelf for a government-run alternative corpus.

What should I know before requesting a sample?

Three things. First, the response shape is documented here from observed structure - the nf_ convention is consistent across client libraries and live responses, but the full attribute list is pinned against live records when your sample is cut, not before. Second, geography has a center of gravity: North American barcodes resolve best, so test any non-US or non-Canadian private-label items in your sample rather than assuming parity. Third, nutrient completeness varies by item - potassium appears only where labels carry it, and restaurant items report per menu serving rather than per 100 grams. Name the brands, chains and cuts you care about, and the sample returns in exactly the schema shown above.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - Nutritionix Natural Language & Grocery API (one row per food item)
fieldtypedefinitionexample
food_namestringName of the matched food item returned by natural-language parsing, autocomplete search and item lookup alike.Wheat Bread
brand_namestringBrand or restaurant chain associated with the item; splits 48,317 grocery brands from 860 monitored chains.Nature's Own
serving_qtynumberQuantity of the default serving unit; pairs with serving_unit to define what the nutrient panel describes.1
serving_unitstringUnit of the default serving - g, oz, cup, slice, sandwich - the normalization basis for cross-product comparison.slice
serving_weight_gramsnumberGram weight of the parsed serving, converting household measures to a mass basis in one step.43
nf_caloriesnumberCalories per stated serving.110
nf_total_fatnumberTotal fat in grams per stated serving.1.5
nf_saturated_fatnumberSaturated fat in grams per stated serving.0.0
nf_cholesterolnumberCholesterol in milligrams per stated serving.0.0
nf_sodiumnumberSodium in milligrams per stated serving.170
nf_total_carbohydratenumberTotal carbohydrate in grams per stated serving.20.0
nf_dietary_fibernumberDietary fiber in grams per stated serving.3.0
nf_sugarsnumberSugars in grams per stated serving.2.0
nf_proteinnumberProtein in grams per stated serving.4.0
nf_potassiumnumberPotassium in milligrams per stated serving, present where the label carries it.140
nix_item_idstringPlatform-internal item identifier, stable across renames and reformulations; the exact-item lookup key alongside UPC.513fc9e73fe3ffd4030010af
Additional fields on request-Full nutrient vector with derivation attributes, grocery brand taxonomy with derived wellness claims, multilingual item names across 15+ languages, and the common-foods phrase table (10,477 tags / 38,828 phrases).-

Nutritionix Natural Language & Grocery API - product specification

AttributeValue
IndustryOther Specialty Retail
Records1,266,570 food items: 1,053,256 grocery items across 48,317 brands; 202,837 restaurant items from 860 chains; 10,477 common-food tags mapped to 38,828 phrases
Fields16 documented spine fields, with fuller nutrient vectors, taxonomy and multilingual names folded out on request
Geographic coverageUnited States and Canada primary depth (>92% stated UPC match); international common foods and 15+ language add-ons
Temporal coverageContinuously maintained current-state shelf; 3,000+ grocery items added or updated monthly; 860 chains watched for menu changes
GranularityOne record per food item with per-serving nutrient values, resolvable by UPC, item identifier or free-text phrase
Quality score6/10 on Datadory's rubric, against a catalog average of 7.81 across 1,744 datasets
Delivery cadenceDaily, weekly, or hourly

What teams do with it

  • Barcode-scanned product enrichment Join the 1,053,256-item grocery layer onto your own catalog by UPC so every scanned SKU arrives with a complete per-serving nutrient panel instead of a blank label.
  • Food logging and meal tracking Natural-language parsing turns 'oatmeal with banana' straight into structured rows, which is the mechanic behind roughly 250 million queries a month from consumer health apps.
  • Better-for-you assortment screening Cut the grocery shelf on sugar, sodium and fiber thresholds per serving to build low-sugar, high-protein or low-sodium plan-compliant product lists at category scale.
  • Menu and formulation change monitoring Restaurant coverage across 860 chains means chain menu nutrition shifts are tracked continuously - competitive intelligence that never requires re-transcribing a menu board.
  • Health-app feature building Autocomplete search plus id-stable item identifiers give developers the lookup spine for calorie counters, macro trackers and recipe analyzers without maintaining their own food database.
  • Nutrition research at population scale Researchers profile branded and restaurant foods against dietitian-curated reference items instead of hand-keying labels, with the common-foods layer bridging survey language to packaged reality.

Questions buyers ask

What does the nutritionix natural language grocery api data contain?

Item-level nutrition records for 1,266,570 foods: 1,053,256 barcoded grocery items across 48,317 brands, 202,837 restaurant items from 860 chains, and a dietitian-built common-foods layer mapping 10,477 tags onto 38,828 everyday phrases. Each record carries a per-serving nutrient panel - calories, fat, saturated fat, cholesterol, sodium, carbohydrate, fiber, sugars, protein and potassium - plus brand, serving definition and stable identifiers.

How does natural-language food parsing work on this dataset?

You send a phrase such as 'two scrambled eggs and toast' and receive one structured row per matched food, each with its own name, serving quantity, unit and nutrient panel. A companion common-foods layer - 10,477 dietitian-maintained tags mapped to 38,828 phrases - resolves everyday wording before falling back to branded and restaurant items, which is what keeps conversational logging accurate.

Can I look up products by UPC barcode?

Yes - that is the grocery corpus's core strength. Barcoded items can be resolved directly by UPC, and the stated match rate for US and Canadian scans exceeds 92%. Each hit returns the full item record including the internal item identifier, so a single scan gives you both the nutrient panel and the stable key for future joins.

Does coverage include restaurant menu items as well as groceries?

It does: 202,837 items from 860 chains, held in the same schema as the grocery records and monitored for menu changes, so when a chain reformulates or adds an item the record moves with it. That makes the corpus usable for both retail-shelf enrichment and dining-out tracking from one dictionary.

How current is the data, and can I set my own refresh cadence?

The corpus is continuously maintained - more than 3,000 grocery items are added or updated monthly and restaurant menus are watched for changes. Through Datadory the cadence is yours: API, files, or your warehouse, daily, weekly, or hourly, with each pull validated against the dictionary so upstream edits surface as visible diffs rather than silent surprises.

Can I evaluate the fields and row shapes before committing?

Yes. Name the brands, chains, barcodes or phrase-parsing examples you need and Datadory returns a sample shaped exactly like the dictionary above, cut against live records. Column naming locks at that stage, and any additional fields - full nutrient vectors, brand taxonomy, multilingual names - are confirmed in the sample rather than promised after.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing