Soft Drinks & Non-Alcoholic Beverages

UK data.gov.uk - Soft Drink Datasets

Datadory delivers uk data gov uk soft drink datasets data covering the National Data Library's soft drink slice: 449 catalogue records across 23 result pages, drawing on ONS drinking-behaviour surveys, NHS youth consumption series, GLA food-and-drink retail registers and Defra and FSA drinking-water releases. Delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
United Kingdom - England, Scotland, Wales and Northern Ireland publishers in one result set
How far back
Records run from the 2010s onward with ongoing updates; each record keeps its publishing body's own clock
How fine
Catalogue-level index, one row per published dataset; each child dataset carries its own grain

What is UK data.gov.uk - Soft Drink Datasets?

A search term aimed at the United Kingdom's national data catalogue, and everything relevance matching brings back. Querying 'soft drink' returns 449 catalogue records across 23 result pages, drawn from a catalogue holding tens of thousands of UK public datasets - central departments, agencies, devolved administrations and individual councils among them.

The spread is wider than the label suggests. Drinking-water protection zones, alcohol-consumption statistics and food-and-drink retail locations outnumber conventional beverage market data. The most beverage-relevant anchors on the early pages are the Office for National Statistics' Drinking: adult's behaviour and knowledge survey, NHS Digital's Smoking, drinking and drug use among young people in England series, the Greater London Authority's Food and non alcoholic drinks register, Calderdale Metropolitan Borough Council's Retail Food and Drink list, and drinking-water quality releases from Defra and the Food Standards Agency.

Every record describes itself against one shared pattern - title, publishing organisation, last-updated date, declared file format - which is what turns 23 pages of mixed-government publishing into a single typed, loadable table.

Get a sample of this dataset and we cut the 449 down to the slice you actually need before anything recurring is switched on.

What do the sample rows look like?

Five records from the result set, exactly as the catalogue presents them:

# five of 449 matching catalogue records - fields per the dictionary below
title=Drinking: adult's behaviour and knowledge              publisher=Office for National Statistics
title=Smoking, drinking and drug use among young people
      in England                                             publisher=NHS Digital
title=Food and non alcoholic drinks                          publisher=Greater London Authority
title=Retail Food and Drink                                  publisher=Calderdale Metropolitan Borough Council
title=Aggregates Levy Bulletin                               publisher=HM Revenue and Customs

Read the rows as a relevance map rather than a highlight reel. Two are consumption surveys - adults and young people covered separately, which is uncommon symmetry for demand-side work. Two are retail geography: a capital-wide food-and-drink premises register and a single metropolitan borough's retail list. One is pure noise - a tax bulletin with no beverage connection that the matcher swept in anyway. That hit-rate spread is the honest cost of keyword matching against government publishing, and filtering it out becomes a one-column exercise once the whole set arrives typed.

Which fields does the dataset dictionary define?

Four fields carry every delivered row, defined against the catalogue's own result cards during the August 2026 research pass rather than read off a published codebook - so the dictionary ships at inferred confidence, and every definition is re-checked against live records before your sample leaves.

title and publisher do the identification work: one free-text name per registered dataset and one attributable organisation behind it, from Whitehall departments to single councils. updated is the freshness screen - in a slice where an ONS survey and a borough retail list keep different clocks, sorting on it separates the living from the dormant before anyone builds on either. format declares what each record resolves to, and the vocabulary runs wide: twelve distinct declared formats span the set, from tabular CSV, XLSX and JSON through geographic GeoJSON, SHP, KML, WMS and WFS to ZIP archives and XML, HTML and PDF documents.

Additional fields on request: each record also carries a publisher-attached usage designation, resource-level pointers to the underlying files, topic classifications and harvest lineage. Those fold rather than show because their values vary by publisher and are confirmed against live rows when your sample is prepared - along with an honest account of which of the 449 records publish them in full.

Where does coverage run across geography, time and granularity?

  • Geography: the United Kingdom at all four nations at once. England dominates by volume, but Scottish, Welsh and Northern Irish publishers sit inside the same result set - which is precisely what a single-agency feed cannot give you.
  • Temporal: deliberately mixed. Records run from the 2010s onward with ongoing updates, and each keeps its publishing body's own rhythm - a recurring health survey lands in waves while a premises register trickles. The updated column states each record's case, so staleness is a filterable field rather than a surprise.
  • Granularity: one row per catalogue record. This is a discovery and inventory layer; it counts datasets, not litres or outlets. Once a record earns a slot in your workflow, the underlying dataset's own grain takes over.

Set against the wider Datadory catalog - a mean quality score of 7.81 across all 1,744 datasets - this slice scores 6/10, a band shared by 230 records. Breadth and attribution are the strengths; the discount is for indexing rather than measuring. Pair it with an observation series such as USDA ERS Sugar and Sweeteners Yearbook Tables when your model needs numbers instead of pointers.

How is the data delivered through Datadory?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel your stack already speaks and set the cadence to match the decision being fed - nightly flat files for a research bench, a pipe straight into Snowflake, BigQuery or Redshift, or lookup calls for anything user-facing. Changing frequency afterwards is a settings conversation, not a re-integration project.

Every delivery ships the four-field dictionary above unchanged, with sample rows attached for validation. Name the publishers, themes or formats you care about - consumption surveys only, or retail registers for named councils - and the sample comes back shaped to that slice before any commitment.

Who uses this data, and for what?

  • Category and competitive teams watch the public-health perimeter around soft drinks: fresh ONS and NHS consumption waves signal where demand for low- and no-sugar lines moves next.
  • Retail location and field-sales planning - the Greater London Authority's food-and-drink premises register and council retail lists turn outlet mapping into a join instead of a manual trawl through council websites.
  • Policy and tax researchers establish what the UK actually publishes on sugary-drink taxation - including the documented absence of Soft Drinks Industry Levy statistics from this slice, which is evidence worth holding before commissioning primary collection.
  • Water-sourcing diligence - Defra and FSA drinking-water quality releases support siting and compliance screens for bottling operations.
  • Data scientists and ML engineers sweep 449 consistently-typed titles and publisher names as a labelled corpus for classification and entity-extraction models.
  • Journalists and academics cite named surveys and registers with a clear line back to each publishing body.

Which personas get the most value?

Competitive-intel and product teams treat the consumption surveys as the demand-side early-warning radar - competitive intel product teams use cases. Market researchers and consultants inventory what UK public data already covers before spending on fieldwork - market researchers use cases. Investors and quants read public-health publishing cadence as a category-pressure signal - investors quants use cases. Data scientists and ML engineers inherit one typed schema across 449 heterogeneous publishers - data scientists use cases. Developers and builders wrap products around a catalogue whose rows share one field structure regardless of which body published them - developers builders use cases. Persona-by-persona detail lives on the Soft Drinks & Non-Alcoholic Beverages data hub.

What should I know before requesting a sample?

Three things worth knowing upfront.

First, this is an index, not the datasets at full depth. Each row points at a catalogue record whose own size, cadence and structure differ; tell us which records caught your eye and we scope the deeper extraction honestly.

Second, market data is the minority party here. Water protection zones, alcohol-consumption statistics and food-and-drink retail locations dominate the 449, so the valuable work is isolating the fraction that matters - which is precisely what a scoped sample settles before anything recurring is built.

Third, if you arrived hunting Soft Drinks Industry Levy revenue or registrant figures, they did not surface in this slice: levy-related searches returned other tax bulletins, and the gov.uk SDIL collection holds guidance documents rather than statistical tables. We say so in the sample rather than pad the feed with near-misses. Decide the cadence - daily, weekly, hourly - after the sample validates, not before. Start at request the sample.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - UK data.gov.uk - Soft Drink Datasets (inferred confidence, verified at sampling)
fieldtypedefinitionexample
titlestringDataset title as registered by the publishing authority.Food and non alcoholic drinks
publisherstringPublishing organisation behind the record - central department, agency or local council.Greater London Authority
updateddateDate the catalogue record was last updated; the freshness screen across a slice whose records keep different clocks. Vintages run from the 2010s onward.date-stamped per record
formatstringDeclared format of the resources behind the record - twelve distinct values span the catalogue, from CSV, XLSX and JSON to WMS, WFS, SHP, GeoJSON, KML, ZIP, XML, HTML, PDF and XLS.CSV

Questions buyers ask

How many datasets match 'soft drink' in the UK catalogue?

449 catalogue records across 23 result pages, retrieved from a national catalogue holding tens of thousands of UK public datasets. The count drifts as public bodies publish, update and retire records, and the set spans everything from ONS consumption surveys to single-council retail lists - delivered as one typed table rather than 23 pages of links.

Which publishers contribute the most beverage-relevant records?

The Office for National Statistics (Drinking: adult's behaviour and knowledge), NHS Digital (Smoking, drinking and drug use among young people in England), the Greater London Authority (Food and non alcoholic drinks) and Calderdale Metropolitan Borough Council (Retail Food and Drink) anchor the early pages, with Defra and the Food Standards Agency supplying drinking-water quality releases.

Does this slice contain Soft Drinks Industry Levy statistics?

No - and the negative result is documented. Searches across levy-related terms returned Aggregates Levy and Climate Change Levy bulletins rather than SDIL statistics, even when filtered to HM Revenue and Customs as publisher, and the gov.uk SDIL collection holds guidance documents rather than statistical tables. Scope levy analytics against that reality rather than an assumed feed.

Is this one dataset or many?

Many. Each matching record is a standalone dataset owned by its publishing body - a national survey, a city premises register, a regulator's water-quality release. The catalogue binds them by describing every row against the same four-field pattern, which is what makes the whole slice queryable as one unit.

How current are the records?

Mixed by design. Records run from the 2010s onward with ongoing updates, and each keeps its publishing body's own clock - a recurring health survey and a dormant council list can sit pages apart. The updated date ships as a first-class field, so recency screening happens before analysis rather than after a surprise.

Can a sample be scoped to particular publishers or themes?

Yes. Name the publishers, themes or formats - consumption surveys only, retail registers for named councils, water-quality releases - and Datadory returns rows shaped exactly like the dictionary above, with the folded fields confirmed against live records. Samples precede any commitment, and the sample schema is the shipped schema.

Notes on this record

  • A search slice, not a curated register 449 hits are what relevance matching returns, not an editorial selection of soft-drinks data. The boundary moves as public bodies publish, retire and rename records.
  • Market data is the minority party Drinking-water protection zones, alcohol-consumption statistics and food-and-drink retail locations outnumber conventional beverage market data in the result set. Separating the fraction that matters is part of the job, not an afterthought.
  • The levy absence is itself a finding No HMRC Soft Drinks Industry Levy statistics surfaced in the catalogue this session, and the gov.uk SDIL collection holds guidance documents only. Knowing what a government does not publish is worth as much as knowing what it does.
  • Twelve declared formats travel Declared formats across the catalogue span CSV, XLSX, JSON, XML, HTML, PDF, ZIP, GeoJSON, KML, SHP, WMS and WFS - spreadsheets, map services and documents side by side under one dictionary.
  • Scored honestly It scores 6/10 against a Datadory catalog mean of 7.81 across 1,744 datasets: breadth and attribution are strong, while depth lives inside each child dataset. Pair it with a measurement series such as CDC NHANES intake data when you need observations.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing