Datadory notebook
Is there a free API for tobacco statistics? The records behind the answer
Datadory delivers tobacco industry data covering every layer behind the tobacco statistics question: a 6,625-row prevalence panel across roughly 265 economies from 2000 to 2024, roughly 300 WHO indicators with MPOWER policy scores for every member state since 2008, smoking-attributable deaths and DALYs for 204 countries plus 660 subnational locations back to 1990, and monthly US removals volumes from calendar 2012 through the July 2026 posting - typed rows, field dictionaries attached, delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
Which datasets carry tobacco statistics?
**World Bank API - Tobacco Smoking Indicators (SH.PRV.SMOK*)** shares the top of the ranking with its bulk twin, World Bank - Prevalence of Current Tobacco Use (% of adults), and both hold the slice's only perfect quality scores: 10 out of 10. One age-standardized percentage per economy per year - adults 15+ using any tobacco product, smoked or smokeless, e-cigarettes excluded - across roughly 265 economies and aggregate regions annually from 2000 to 2024, about 6,625 observed rows. Male (SH.PRV.SMOK.MA) and female (SH.PRV.SMOK.FE) siblings add around 5,000 observations apiece on identical keys.
WHO Global Health Observatory Legacy Data Repository - Tobacco widens the lens to roughly 300 tobacco indicators across 194-plus member states: modeled prevalence of smoking, smokeless and e-cigarette use beside the MPOWER policy family, cessation measures and tobacco-tax indicators, at country x year x sex grain. Scored 8. WHO Report on the Global Tobacco Epidemic 2023 (scored 9) carries the policy story in ninth-edition form - every member state graded on all six demand-reduction measures back to 2008, backed by roughly 369,000 structured rows across 29 data files.
IHME GBD Results Tool - Tobacco/Smoking Query Interface owns the outcome layer: deaths, YLLs, YLDs, DALYs, population attributable fractions and summary exposure values for smoking and secondhand smoke, resolved for 204 countries and territories plus 660 subnational locations every year from 1990 to 2023, each estimate carrying its 95% uncertainty interval. Scored 9. Its sibling, IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence (scored 9), compresses the older vintage into six compact files - prevalence, smoker counts and cigarette-equivalents per capita for 1990-2019.
Three more fill specific jobs. Eurostat - Daily Smokers of Cigarettes (sdg_03_30) observes European behavior directly: the share of people aged 15+ smoking boxed cigarettes, cigars, cigarillos or a pipe - vapes and snuff deliberately excluded - by sex across 35 European entities on seven Eurobarometer waves, 2006 to 2023, roughly 620 observations. US TTB - Tobacco Statistics counts what left bond: monthly and annual removals, production, exports, inventories and industry-member counts from calendar 2012 through the July 2026 posting. Our World in Data - Smoking curates about 38 long-run indicators into citation-ready country-year tables. Machine-learning work draws on the Hugging Face tobacco hub - 118 repositories led by CDC STATE System policy tables spanning 1995 to 2024 and an approximately 800,000-row smokefree indoor air corpus.
What does a tobacco statistics row look like?
One genuine record shape out of each family, flattened for reading:
# World Bank SH.PRV.SMOK - one observation = economy x year, adults 15+
country.value : Africa Eastern and Southern countryiso3code : AFE
date : 2023
value : 11.8997926098489 # 11.9% of adults use tobacco
# sex splits ride sibling indicator codes on identical keys
SH.PRV.SMOK.MA AFE 2023 19.96 # adult men
SH.PRV.SMOK.FE AFE 2023 3.78 # adult women
# WHO GHO - one row = indicator x country x sex x year, interval attached
indicator : M_Est_tob_curr country : MYS sex : SEX_MLE year : 2022
value : 34.5 interval : 22.0 - 47.0
# MPOWER policy row - same schema, integer achievement level
IndicatorCode : TOBACCO_MPOWER_OVERVIEW SpatialDim : SDN TimeDim : 2022
Value : 4 # highest level on that measure
# US TTB - one observation = period x group x basis x product line
month : 1 year : 2012 group : Removed Taxable detail : Category Total
value : 22376933943 # 22.4 billion sticks count_IMs : 106Three design choices reward attention. First, identity travels with the number everywhere: ISO-3 codes sit beside display names in the prevalence rows, and indicator, country, sex and year resolve in the same WHO row, so panels assemble with a join rather than a parsing job. Second, honesty is a column, not a footnote - Malaysia's 34.5% male current-tobacco-use estimate for 2022 ships as 34.5 within a 22.0-to-47.0 band, and TTB stamps a reporter count on every value. Third, the sex splits argue their own point: in the African Eastern and Southern sample, 11.9% of adults used tobacco while men read 19.96% against women at 3.78% - an asymmetry any excise or volume forecast has to respect.
Which tobacco question lands on which record?
The pool divides by job-to-be-done rather than by publisher, so choosing the right record first saves real engineering time. As cataloged by Datadory, August 2026:
How deep does tobacco coverage run?
Four clocks run at once, and knowing which one you are on prevents most bad joins.
- Behavioral prevalence runs annually from 2000 through 2024 across roughly 265 economies, though per-economy windows vary with national survey practice, so aligning start years is part of any serious comparison.
- Modeled burden runs every single year from 1990 to 2023 - thirty-four consecutive points per cell rather than intermittent waves, which is what makes curve-fitting honest.
- Policy scores reach back to 2008 per measure, moving in reporting rounds rather than continuous years.
- European surveys move on seven discrete Eurobarometer waves between 2006 and 2023 - intermittent by design, because observations exist only where the survey ran. American volumes run monthly from calendar 2012 through the July 2026 posting, restated in place so prior periods reflect current vintages rather than frozen first prints.
Delivery smooths the staircase. Most teams take historical depth once as a backfill, then keep new periods rotating in on whatever rhythm their models expect - daily, weekly, or hourly, your call - instead of polling each vintage themselves.
Which traps distort tobacco statistics?
Mixing measures. A Eurostat self-reported share, a TTB removals count and a GBD attributable-death total are three different claims about the same market. Joined into one chart without labeling the units, they draw a series that looks continuous and is not.
Ignoring the interval. Modeled estimates ship their uncertainty bands attached - Cameroon's 9.1% reading for 2023 sits inside a tight 6.5-to-11.8 range while Malaysia's spans twenty-five points. Dropping the band manufactures precision nobody verified.
Forgetting who smokes. Male prevalence sits far above female nearly everywhere - 19.96% against 3.78% in the sample rows above - so a total-population average quietly folds demographic composition into what reads as a behavior change.
Who builds on tobacco statistics?
Market researchers and consultants size nicotine demand and benchmark markets on one harmonized definition instead of stitching national statistics offices together - see market researchers use cases.
Investors and quant researchers read long-run morbidity and mortality as demand-side fundamentals for insurers, treatment providers and pharma, with subnational detail for prioritization.
Excise and public-health analysts set modeled prevalence beside administrative removals volumes, treating the gap between the two as the interesting object rather than noise.
Data scientists and ML teams train document classifiers on labeled corpora - the 3,482-image tobacco3482 set leads the Hugging Face shelf - while journalists and academics cite burden figures that ship with their own confidence bounds rather than bare headlines.
All six jobs run off the same typed rows, delivered daily, weekly, or hourly - your call.
Why get tobacco statistics through Datadory?
Because the numbers were never the hard part - shape is. Four record families with different key structures. Uncertainty columns that belong in the loss function rather than a methods appendix. Survey-wave gaps that read as holes until someone checks the design. Reporter counts that decide whether a product line is a market or a duopoly. Reconciling all of that eats days before any analysis starts.
Request a sample: name the countries, sexes, years and measures, and real rows come back cut to them before anything recurring starts.
Where to go next
This page is one thread of the tobacco stack. Open the tobacco data hub for the complete scored index of the industry slice, or the ranked best tobacco datasets shortlist for the top of it. Smoking prevalence by country walks the prevalence panels record by record, cigarette consumption per capita covers the converted demand measure, and US cigarette production statistics digs into the American supply side. Dataset pages document the World Bank indicator family, the WHO Global Health Observatory repository and US TTB volumes down to the field dictionary.
| Question you are asking | Record to reach for | Why it wins |
|---|---|---|
| National trend, country ranking or regional rollup? | World Bank API - Tobacco Smoking Indicators (SH.PRV.SMOK*) | One harmonized percentage per economy-year for ~265 economies, 2000-2024, with pre-computed aggregates beside individual countries |
| Splitting behavior by sex? | World Bank prevalence panel plus .MA and .FE splits | ~16,600 combined rows on identical economy-year keys - no reshaping to compare men and women |
| Measuring European smoking behavior directly? | Eurostat - Daily Smokers of Cigarettes (sdg_03_30) | Self-reported daily smoking by sex across 35 European entities on seven survey waves, 2006-2023, combustibles only |
| Tracking US market throughput and excise volumes? | US TTB - Tobacco Statistics | Monthly removals, production, exports, inventories and permittee counts in sticks and pounds, 2012 through the July 2026 posting |
| Need a citable long-run chart quickly? | Our World in Data - Smoking | About 38 curated indicators from roughly 1990 into the 2020s, cleaned to country-year tables |
| Training a classifier or mining policy text? | Hugging Face Datasets - Tobacco Hub (118 datasets) | ML-ready corpora from CDC policy tables to an ~800,000-row smokefree indoor air set and the tobacco3482 image collection |
| Dimension | Coverage |
|---|---|
| Geographic | ~265 economies and aggregates in the prevalence panel; 194+ WHO member states in the indicator repository; 204 countries plus 660 subnational locations in modeled burden; 35 European entities in the survey series; United States national in the volumes ledger |
| Temporal | Annual 2000-2024 (prevalence); every year 1990-2023 (burden); measure rounds since 2008 (MPOWER); survey waves 2006-2023 (Europe); monthly and annual 2012 through the July 2026 posting (US volumes) |
| Granularity | Economy x year with sex-split siblings; country x year x sex x indicator; location x age group x sex x year x cause-or-risk; month-or-year x statistical group x measurement basis x product line |
| Measure families | Behavioral prevalence, modeled burden, policy achievement scores, market volumes - plus curated long-run indicators and ML corpora |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
World Bank - Prevalence of Current Tobacco Use (% of adults)
countryiso3code · date · value …+3 more
World Bank API - Tobacco Smoking Indicators (SH.PRV.SMOK*)
WHO Global Health Observatory Legacy Data Repository - Tobacco
WHO Report on the Global Tobacco Epidemic 2023
IHME GBD Results Tool - Tobacco/Smoking Query Interface
IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019
location_name · sex · age_group_name …+4 more
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Is there a free API for tobacco statistics?
The question hides two decisions: what exists, and how it arrives. On existence, the field is deep - a 6,625-row prevalence panel spanning roughly 265 economies, roughly 300 WHO indicators with MPOWER policy scores, modeled deaths and DALYs for 204 countries plus 660 subnational locations, and monthly US removals volumes. On arrival, Datadory delivers all of it as typed rows by API, scheduled files, or straight into your warehouse - daily, weekly, or hourly - so transport stops being the constraint.
Which dataset measures smoking prevalence by country?
World Bank indicator SH.PRV.SMOK: the percentage of adults aged 15 and over currently using any tobacco product, age-standardized, one observation per economy per year for roughly 265 economies and aggregates annually from 2000 to 2024 - about 6,625 rows, with male and female splits adding around 5,000 observations apiece on identical keys. It is the only perfect-score record in the tobacco slice.
Where do smoking-attributable deaths and DALYs come from?
IHME's Global Burden of Disease program, queried through the GBD Results Tool tobacco interface. It models deaths, YLLs, YLDs, DALYs, attributable fractions and summary exposure values for smoking and secondhand smoke across 204 countries and territories plus 660 subnational locations, every year from 1990 to 2023, with a 95% uncertainty interval traveling on every estimate.
Which data tracks tobacco control policy by country?
The MPOWER family: every WHO member state scored 1-to-4 on each of the six demand-reduction measures - monitoring, smoke-free places, cessation support, warnings, advertising bans and taxation - back to 2008, published in the Report on the Global Tobacco Epidemic editions and mirrored in the Global Health Observatory alongside roughly 300 other tobacco indicators.
What US tobacco volume data exists?
US TTB - Tobacco Statistics, the federal ledger the excise system bills against: monthly and annual taxable removals, manufacturing volumes, tax-exempt exports, close-of-business inventories and industry-member counts for cigarettes, cigars, smokeless and pipe tobacco, in stick counts and pounds, running from calendar 2012 through the July 2026 posting. January 2012 alone logged 22.4 billion sticks removed, filed by 106 permittees.