Soft Drinks & Non-Alcoholic Beverages · UN Food and Agriculture Organization (FAO)
FAOSTAT Food and Agriculture Data
Datadory delivers faostat food and agriculture data data covering the agricultural inputs behind every bottled drink - sugar cane, sugar beet, oranges and other citrus, tea leaves, coffee green and hundreds of their neighbors - for over 245 countries and territories from 1961 to the most recent year available: production, harvested area and yield beside bilateral import and export flows, food balance sheets and producer price indexes, one tidy observation per row with its provenance flag attached. Delivered daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
What is FAOSTAT Food and Agriculture Data?
FAOSTAT Food and Agriculture Data is the Soft Drinks & Non-Alcoholic Beverages catalog's record over the UN Food and Agriculture Organization's core statistical database: the account of what the world grows, trades, processes and prices, kept by more than 245 countries and territories and reaching back to 1961.
For a beverage business the relevant shelf is specific. Crop production holds the sweetener and flavor inputs by name - sugar cane (item code 156), sugar beet (157), oranges (490), coffee green (656), tea leaves (667) - each carrying three elements per country-year: area harvested in hectares, production in tonnes, yield in hectograms per hectare. The detailed trade matrix turns those same items into bilateral flows, one row per reporter-partner-item-element-year, so an import dependency question is a filter rather than a research project. Food balance sheets add the supply-versus-use ledger per country, and producer price indexes date what the raw material cost at the farm gate.
Two design choices make the whole thing usable at scale. Every domain ships the identical row anatomy - area, item, element, year, unit, value, flag - so skills learned on one transfer to all of them. And every figure declares its own grade: officially reported, estimated, imputed by a receiving agency, contributed by an external organization, or structurally missing. Get a sample of this dataset cut to the countries, items and horizon your models actually touch.
What do the sample rows look like?
One observation per row, fully labeled, identifiers attached. A captured row exactly as it ships, above the five item codes every soft-drink ingredient model names first:
# captured observation -- December 2025 build
Area Code : 5817 Area : Net Food Importing Developing Countries (NFIDCs)
Item : Vegetables Primary Element: Production
Year : 2024 Unit : t Flag : A (official figure)
# the beverage-input shelf -- measure slots fill when your sample is cut
156 Sugar cane Area harvested | Production | Yield ha | t | hg/ha
157 Sugar beet same elements, plus trade quantity/value rows
490 Oranges same spine
656 Coffee green same spine
667 Tea leaves same spineThe first row is doing quiet work: area code 5817 is not a country but a regional aggregate - the Net Food Importing Developing Countries grouping - living in the same Area column as France or Kenya. That is the schema's one sharp edge. Sum without excluding aggregates and a world total double-counts itself. The five item codes below it are the point: swap the area filter and the identical query returns Brazilian cane, Spanish oranges or Kenyan tea, because the layout never moves. Values print at full precision - six decimals in normalized extracts - and every row carries its flag, which is why a government-reported figure can be told apart from an agency's estimate after delivery rather than during an argument.
What fields does the dataset include?
Twelve documented fields define every observation, each definition checked against the record during research rather than inferred from prose. They stack into a natural key - area x item x element x year - which means joins against your own commodity master or sourcing map hold up without fuzzy matching, and the same query pattern serves cane tonnage and tea yields alike.
Three columns deserve a second look before you build. Area mixes countries with FAO aggregates and special groupings, so exclusion lists belong in every sum. Value prints at full published precision - six decimals in the normalized extracts - and rewards deliberate casting. And Flag is the column analysts under-use: it is the difference between what a country declared and what the receiving agency filled in, and reading it before trend work keeps reconstruction visible instead of silent.
Domain extensions fold under additional fields on request: reporter-partner columns from the trade matrix, the balance-sheet element sets, producer price series and the decode tables that turn codes into words. Name them when you request the sample and they arrive attached to the core rows rather than left as folders to reconcile.
What does coverage look like across geography, time and granularity?
Geography - over 245 countries and territories, plus FAO regional groupings and special aggregates carried beside them. Country is the analytical grain; regional rows are conveniences built on top, not independent measurements. On the trade side the matrix distinguishes roughly 232 reporting from around 255 partner areas, so a destination-market view can include suppliers that never report directly.
Temporal - depth varies deliberately by domain. Crop production runs from 1961 through 2024 in the build reviewed. The detailed trade matrix starts in the mid-1980s. Producer prices begin annually in 1991 and go monthly from January 2010. Balance sheets split into pre-2010 and current-methodology editions, which matters if you span the boundary. Nothing pretends to a rhythm it does not have.
Granularity - one observation per area x item x element x year everywhere on the production side; one row per reporter x partner x item x element x year in the matrix. Annual throughout, with no interpolation upward and no sub-national resolution anywhere: these are national accounts compiled by an international agency, and anyone promising province-level cane tonnage from this source is selling a different dataset.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
You pick the channel and the cadence; the field dictionary above travels unchanged through all three. These are annual statistics at the core, so a weekly warehouse load keeps dashboards current between builds and nobody has to babysit the pipeline. Every delivery ships the complete twelve-field dictionary, the sample rows and a coverage profile mapped to the countries, items and years you named - plus the reference tables that decode area, item, element and flag codes - so the extract lands pre-cut rather than as a pile for your team to sort. The sample comes first either way.
Who uses this data, and for what?
A sixty-plus-year supply ledger earns its keep in five specific jobs:
- Ingredient supply and cost modeling - production, yield and harvested-area series for sugar crops, citrus, tea and coffee give procurement and FP&A an official benchmark for the growing side of input costs.
- Sourcing-risk screens - production beside bilateral trade isolates concentration: one country's share of a destination's orange imports, one region's grip on cane exportable surplus.
- Category demand context - food balance sheets separate what a market produces from what it has available to consume, the distinction most category models quietly skip.
- Input-price feature engineering - producer price indexes, monthly since 2010, join onto commodity futures and CPI work under the same area-item keys.
- ESG and origin claims - officially flagged country series give disclosure and marketing claims citation-grade provenance, with the flag column marking the estimates.
Which personas get the most value?
Market researchers and consultants (relevance 3/3) size ingredient markets across 245+ territories from one consistent schema; see market researchers use cases. Data scientists and ML engineers (3/3) get six-decade country-commodity panels under a twelve-field spine with flags kept as features; see data scientists use cases. Procurement and supply-chain teams (3/3) watch the growing regions behind every ingredient before contracts are negotiated. Journalists, academics and students (3/3) cite the UN's standard agricultural statistics with provenance attached to every figure. Developers building data products (2/3) wire one uniform row shape into dashboards across every domain; see developers builders use cases. Investors and quant researchers (2/3) read crop supply momentum as an exogenous signal on beverage-input cost cycles; see investors quants use cases. Sales teams sit near zero relevance: national crop tonnage carries no contact-level or firmographic signal.
How does it compare to alternatives in its slice?
Within soft drinks and non-alcoholic beverages data, this record owns the global supply side: what the world grows, trades and balances for the inputs a drink is made of. The neighbors own different jobs. USDA ERS Sugar and Sweeteners Yearbook Tables goes deep on American sweetener economics - prices, deliveries, tariff-rate quotas - where this set goes broad but stops at the farm gate. CDC NHANES - Dietary Intake & Beverage Consumption Data measures who actually drank what, person by person, survey-style - demand-side microdata this administrative compilation cannot approximate. BLS Consumer Price Data Tools (Nonalcoholic Beverages CPI) tracks what consumers pay at the register rather than what the raw material cost at origin. The head-to-head with NHANES is laid out in vs CDC NHANES, and the deeper FAO cuts live on sibling records: FAOSTAT - Crops and Livestock Products for the production universe at full depth and FAOSTAT - Detailed Trade Matrix for the bilateral flows. If the question is "where does the world's sweetener, citrus, tea and coffee supply come from, since 1961", this is the set that answers it.
What should I know before requesting a sample?
Four things worth knowing upfront.
First, aggregates live beside countries. Area code 5817 is a grouping, not a geography you can ship from, and it sits in the same column as every nation. Build exclusion lists into ingestion on day one or world totals will quietly double-count.
Second, precision is real but fragile. Values print with six decimals in the normalized extracts; naive integer casting corrupts tonnages before you ever see the error. Cast once, centrally, and keep the raw strings if audit matters.
Third, domains have different clocks and edges. Production reaches furthest back, trade starts later, prices go monthly only in the modern era, and balance-sheet methodology splits old from new. Panels that cross those boundaries need explicit rules rather than optimistic joins.
Fourth, the field list was verified from the record rather than exercised end-to-end on live pulls, which is why the domain extensions sit folded under additional fields on request. Say which countries, items and years matter and the sample comes back shaped to exactly that scope, complete dictionary attached.
Which datasets pair well with this one?
Notes that pair well with this page:
- USDA ERS Sugar and Sweeteners Yearbook Tables - the American sweetener ledger; global breadth beside national depth covers both ends of the sourcing question.
- FAOSTAT - Detailed Trade Matrix - the bilateral cut of the same commodity vocabulary; production tells you who grows, the matrix tells you who buys.
- FAOSTAT - Fertilizers by Nutrient - the upstream input to the inputs, in nutrient tonnes, same country-year grain.
- Glossary primers - food balance sheets, area harvested, production and yield and the FAOSTAT data quality flag explain the dimensions this page leans on.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
Area Code | integer | Numeric FAO identifier for the country, territory or regional grouping the observation belongs to. | 5817 |
Area Code (M49) | string | UN M49 geographic code for the same area, prefixed with an apostrophe in normalized files, so joins to any M49-keyed master hold without a crosswalk. | '902 |
Area | string | Name of the country, territory or regional grouping. Over 245 areas report in, and FAO aggregates such as the NFIDCs sit in the same column as single countries - exclude them before summing. | Net Food Importing Developing Countries (NFIDCs) |
Item Code | integer | FAO commodity identifier for the crop, livestock product or processed good measured. | 156 |
Item Code (CPC) | string | UN Central Product Classification code mapped onto the item, keeping commodity identity joinable outside the FAO vocabulary. | 01110 |
Item | string | Commodity name. The beverage-input shelf reads sugar cane, sugar beet, oranges and other citrus, tea leaves and coffee green, beside several hundred further crops and products. | Sugar cane |
Element Code | integer | Identifier for what is being measured about the item: area harvested (5312), production (5510) and yield (5610) on the production side, import and export quantities and values on the trade side, supply-and-use lines on the balance sheets. | 5510 |
Element | string | Human-readable name of the measure riding on the row. | Production |
Year | integer | Reference year of the observation, annual throughout, running from 1961 to the most recent year published in the build you receive. | 2024 |
Unit | string | Unit attached to the value - tonnes, hectares, kilograms per hectare, thousand US dollars, head - declared per row rather than per table. | t |
Value | number | The reported quantity itself at full published precision. Normalized extracts print six decimal places, so cast deliberately before summing; a naive integer read silently corrupts tonnages. | 74590.000000 |
Flag | enum | Provenance of the figure: A official, E estimated, I imputed by a receiving agency, X supplied by an external organization, M missing value that cannot exist. Filter on it before quoting any country figure. | A |
Coverage at a glance
| Dimension | Value |
|---|---|
| Geography | World - over 245 countries and territories plus FAO regional groupings and special aggregates; the trade matrix distinguishes roughly 232 reporting from some 255 partner areas |
| Temporal | 1961 through 2024 in the production domain reviewed (December 2025 build); trade runs from the mid-1980s, producer prices annually from 1991 and monthly since January 2010, balance-sheet editions split old and new methodology |
| Granularity | One observation per area x item x element x year; one row per reporter x partner x item x element x year in the bilateral matrix - annual throughout, nothing interpolated upward |
What teams do with it
- Ingredient supply and cost modeling Six decades of production, yield and harvested area for sugar crops, citrus, tea and coffee give procurement teams an official benchmark for where input supply comes from and how it moves.
- Sourcing-risk screens Concentration analysis on who grows and who ships what - one country's bad harvest against global supply, or a single exporter's share of a destination's orange inflow - falls out of production and trade rows sharing one key.
- Category demand context Food balance sheets put supply against use per country, so a beverage category model can separate what a market produces from what it actually has available to consume.
- Input-price feature engineering Producer price indexes for crops and processed goods extend cost curves back to 1991 annually and 2010 monthly, ready to join onto commodity futures or CPI work.
- ESG and origin claims Officially flagged country-level production series support origin statements and disclosure work with citation-grade provenance on every number.
Questions buyers ask
What does FAOSTAT Food and Agriculture Data cover?
Country-level agricultural statistics for over 245 countries and territories from 1961 onward: crop and livestock production with area harvested, yield and output; bilateral import and export flows by commodity pair; food balance sheets putting supply against use; and producer price indexes. For soft drinks the headline items are sugar cane, sugar beet, oranges and other citrus, tea leaves and coffee green.
How far back does the data go?
To 1961 on the production side, continuously, in the same row format throughout. The detailed trade matrix starts in the mid-1980s; producer prices run annually from 1991 and monthly from January 2010; balance-sheet editions split into a legacy and a current-methodology series around 2010.
Is the data country-level or can I get sub-national detail?
Country-level throughout. Over 245 countries and territories report, alongside FAO regional groupings and special aggregates that occupy the same Area column and must be excluded before summing. There is no province or state resolution anywhere in the corpus.
What do the flags on each row mean?
Five provenance codes: A official figure, E estimated value, I imputed by a receiving agency, X figure supplied by an external organization, and M a missing value that cannot exist. Filter on Flag before trend work - a decade of imputed values plotted beside officials shows you exactly where the series is solid.
Why do values sometimes look oddly precise?
Normalized extracts print full published precision - six decimal places. That is a feature for reconciliation and a trap for careless casting: read Value as a string or a decimal type, not an integer, and sums stay honest.
Can a sample be scoped to my countries, items and years?
Yes. Name the countries, the item codes - say sugar cane 156 and oranges 490 - and the horizon, and the sample arrives shaped to that scope with the full field dictionary and the reference tables attached. Samples precede any commitment, and the schema you see in the sample is the schema you ship against.
Notes on this record
- Aggregates are not countries FAO regional groupings and special aggregates such as the NFIDCs (area code 5817) share the Area column with individual nations. Exclusion lists belong in ingestion on day one - otherwise every world total double-counts.
- The flag is the evidence grade A marks an officially reported figure, E an estimate, I an agency imputation, X a contribution from an external organization. Reading the flag beside the value separates what a country declared from what was filled in on its behalf.
- One row shape, many domains Production, trade, balance sheets and prices all resolve to area x item x element x year. Skills and pipeline code learned on one domain transfer to the rest - the layout never moves, only the element list does.
- Scored 9/10 in the catalog Datadory scores this record 9 out of 10 against a catalog mean of 7.81 across all 1,744 cataloged datasets - breadth, schema discipline and provenance flags doing the lifting.
Datasets that pair with this one
- FAOSTAT - Detailed Trade Matrix When the question shifts from how much is grown to who sells it to whom, the bilateral matrix carries the same commodity vocabulary one dimension further.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.