Datadory notebook

GBD smoking attributable deaths: modeled burden, delivered as rows

Datadory delivers tobacco data covering GBD smoking attributable deaths end to end: modeled deaths, YLLs, YLDs and DALYs for the Smoking and Secondhand smoke risk factors across 204 countries and 660 subnational locations, every year 1990 through 2023 with 95% uncertainty intervals on every estimate, beside citation-ready chart series reaching the mid-2020s - delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What are GBD smoking attributable deaths?

Not counted one by one. Nobody tallies headstones marked "smoking"; epidemiologists estimate what mortality and disease burden would remain if smoking disappeared, then publish the difference as formal measures. IHME's Global Burden of Disease program is the machinery behind almost every attributable-death figure quoted in public health, and the difference arrives as four related quantities: deaths, DALYs, and their two components - years of life lost (YLLs) and years lived with disability (YLDs). Each ships with a 95% uncertainty interval, because a modeled quantity without its range is a press release wearing a lab coat.

The scale anchor most quotations reach for comes from IHME's own smoking topic record: 1.14 billion smokers in 2019, 7.7 million deaths and 200 million DALYs attributable to smoked tobacco in that reference year, and 82.6% of current smokers reporting they started between ages 14 and 25. Behind those three headlines sits the full cube: estimates spanning 204 countries and territories plus 660 subnational locations, every single year from 1990 through 2023 in the GBD 2023 release.

A vocabulary trap distorts searches before any data lands. Attributable burden covers fatal and non-fatal harm together, so filtering a query to deaths alone understates the total - which is why health-economics teams pull DALYs in the same delivery and reconstruct the fatal share themselves. See Global Burden of Disease for the program behind the estimates.

Which records carry smoking-attributable burden?

Three answer the question directly, and they divide by job rather than by publisher:

  1. IHME GBD Results Tool - Tobacco/Smoking Query Interface - the cube. Deaths, YLLs, YLDs, DALYs, prevalence, summary exposure values and population attributable fractions for the Smoking and Secondhand smoke risk factors, addressable by location, age band, sex, year, measure and metric. It scores 9 on Datadory's rubric, and its subnational depth - 660 locations beneath the 204 national units - is unique in the tobacco pool.
  2. Our World in Data - Smoking - the quotation layer. About 38 chart-backed tables including death rates from smoking and from secondhand smoke, the share of deaths attributed to smoking, lung-cancer deaths and tobacco-attributable cancer deaths, harmonized onto one entity-year frame from providers including IHME's GBD study. Also 9.
  3. HealthData.org - Tobacco Disease & Injury Factsheets Data - the headline shelf. Dozens of topic statistics - the 7.7 million deaths, the 1.14 billion smokers - pinned to the GBD 2019 reference year with their Lancet citations one field away. Narrative-grained rather than row-grained, which caps it at 5.

Beside them sits the exposure side of the ledger: the IHME GHDx GBD 2019 smoking prevalence record packages age-standardized prevalence, smoker counts and cigarette-equivalents per capita for the same 204 countries, deliberately excluding attributable deaths so the two halves stay distinct. Sixteen primary records make up the tobacco pool overall, averaging 8.4 on Datadory's rubric against a catalog-wide mean of 7.81 across 1,744 datasets.

What does one attributable-burden row look like?

One estimate row, addressed down to a single cell of the cube:

# smoking-attributable DALYs - one location x age x sex x year cell
measure        : DALYs            # Deaths, YLLs, YLDs, DALYs, Prevalence,
                                 # Summary Exposure Value, Attribution Fraction...
location_name  : Japan            # 204 national units or 660 subnational ones
sex_name       : Both             # Male / Female / Both
age_group_name : All ages         # standard bands, early neonatal to 95-plus
rei_name       : Smoking          # Secondhand smoke is its own risk factor
metric         : Number           # Number / Rate / Percent
year           : 2019             # continuous annual series, 1990-2023
val            : <point estimate>
lower_val      : <95% UI lower>   # the interval rides the row, always
upper_val      : <95% UI upper>

Three things arrive pre-decided in that shape. The measure is explicit on every row, so "deaths" and "DALYs" never blur into one column. The interval travels beside the point estimate instead of living in a methods appendix, which turns sensitivity analysis into a column operation. And the metric convention is declared - Number, Rate or Percent - so a rate for Japan and a count for Brazil never share an axis by accident. Cause-level rows use the same spine with cause_name replacing rei_name, which lets one delivery hold both "burden attributed to smoking" and "burden from tracheal, bronchus and lung cancer" on identical keys.

Which record answers a deadline question fastest?

Depends who is asking, and the honest rule splits cleanly. Our World in Data when one number with provenance has to be on screen in minutes: pick the chart - share of deaths attributed to smoking, say - read the harmonized entity-year rows, cite the provider named on the face of the series. Point values only, no intervals, and coverage varies by chart, generally running from about 1990 out to 2022-2023. The GBD cube when a reviewer will ask how the attribution was constructed: any age band, any subnational unit, intervals included, one set of assumptions holding the whole extract together.

HealthData.org's topic record serves a third job - the framing sentence. Its statistics stay pinned to the GBD 2019 reference year even as newer rounds publish, which is a feature for briefing decks that need stable citable numbers and a liability for anyone modeling the present. Treat it as citation scaffolding, not as modeling input.

Teams doing sustained work usually keep both postures: a curated chart series as the screenshot, a shaped GBD extract as the audit trail. Deliveries pin the source release per extract, so a figure quoted this quarter reproduces next quarter unless a value genuinely moved - revisions land as identified rounds, never as silent overwrites.

Can attributable burden be joined to exposure and policy data?

Burden gains meaning beside behavior and policy, and the pool supplies all three layers on compatible keys.

On exposure, the World Bank prevalence panel reports the share of adults using any tobacco product for roughly 265 economies annually from 2000 to 2024 - 6,625 observations on the headline series, age-standardized to the WHO Standard Population, with male and female variants near 5,000 rows apiece riding sibling codes on identical economy-year keys. On policy, the WHO Report on the Global Tobacco Epidemic 2023 grades every member state 1-to-4 on each MPOWER measure back to 2008 - the 2023 edition reported smoke-free legislation covering roughly a third of the world's population - while the WHO Global Health Observatory collection carries about 300 tobacco indicators with bounds embedded in every observation; Malaysia's modeled male prevalence for 2022 arrives as 34.5 within a 22.0-to-47.0 interval. Europe adds a surveyed mirror: Eurostat sdg_03_30 counts daily smokers of cigarettes, cigars, cigarillos or a pipe by sex across 35 entities in seven survey waves from 2006 to 2023.

Three caveats keep such joins honest. These are different estimators - modeled intervals, self-reported survey shares, policy grades - so splicing their levels into one series draws a line that does not exist; compare them, label them, cite them separately. Regional composites such as EU27_2020 or Africa Eastern and Southern ride the same schema as sovereign states, and a ranking that fails to filter them ranks Germany against an entire continent. And Eurostat's timeline is a wave grid, not a clock - 2006, 2009, 2012, 2014, 2017, 2020, 2023 - so interpolating between observed years is a modeling decision to declare, not a default to hide. Quoting attributable burden with its interval beside the point estimate is what separates a defensible analysis from a headline.

Who builds on smoking-attributable death data?

Five jobs this evidence settles outright:

  1. Health-economics cases - price attributable DALYs and YLLs by country or by state, intervals feeding sensitivity analysis instead of getting trimmed for convenience.
  2. Insurer and pharma diligence - read long-run morbidity and mortality trajectories as demand-side fundamentals, with subnational granularity where the market-level question demands it.
  3. Citation-grade journalism and scholarship - every quoted figure resolves to location, age band, sex, year, measure and release.
  4. Regulatory and ESG screening - MPOWER grades beside burden trends show where policy intensity moved and whether burden followed.
  5. Modeling targets - a dense, uncertainty-tagged location-age-sex-year panel gives forecasting work ground truth instead of repackaged press numbers.

API, files, or your warehouse. Daily, weekly, or hourly.

Name the geographies, age bands, sexes, years and measures, and the sample lands cut to that scope with the field dictionary attached - typed rows, parsed years, uncertainty columns intact. The historical backfill arrives first; the feed continues at whatever pace the question needs, with the source release pinned per delivery.

Persona fit has edges worth naming: market researchers get market-level burden sizing; journalists and academics get citable figures with intervals attached; quant teams get slow-moving fundamentals for sector screens. What these records do not measure is behavior at the till or volume through the supply chain - pair cigarette consumption per capita and US cigarette production statistics for the demand and supply sides this evidence deliberately leaves alone.

Where to go next

This page is the burden thread of the tobacco stack. Open the tobacco data hub for every record in the industry with coverage and field dictionaries, or the ranked best tobacco datasets shortlist for the top of the pool. The trade-offs run deeper on the dedicated comparisons: Eurostat sdg_03_30 versus the GBD Results Tool sets observed European behavior against modeled global consequences, and GBD 2019 smoking prevalence versus OWID Smoking puts the exposure archive beside the quotation layer.

Practical order of operations: pull the burden spine first - deaths plus DALYs for the Smoking risk factor, all ages beside the working-age bands, your markets plus the benchmarks - then append exposure and policy layers on the shared country keys once the outcome series is clean. When you want rows rather than reading, request a sample cut to the countries, years and measures on your desk this quarter; the schema in the sample is the schema you ship against.

Records that carry smoking-attributable burden (Datadory catalog, as of August 2026)
RecordWhat it carriesCoverage and grainTemporal depthRole in the stack
IHME GBD Results Tool - Tobacco/Smoking Query InterfaceDeaths, YLLs, YLDs, DALYs, prevalence, summary exposure values and attributable fractions for the Smoking and Secondhand smoke risk factors, 95% uncertainty interval on every estimate204 countries and territories plus 660 subnational locations; location x age x sex x year x cause or risk factor x measureEvery year 1990-2023 (GBD 2023 release)The burden spine - full dimensional control beneath the headline numbers
Our World in Data - SmokingAbout 38 chart-backed tables including death rates from smoking and secondhand smoke, share of deaths attributed to smoking, lung-cancer deaths and tobacco-attributable cancer deathsAll countries worldwide plus world and region aggregates; entity x year x indicatorRoughly 1990 out to 2022-2023, varies by chartThe shortest path from question to quotable number
HealthData.org - Tobacco Disease & Injury Factsheets DataDozens of headline smoking statistics - 1.14 billion smokers, 7.7 million attributable deaths, 200 million attributable DALYs in 2019 - with citations attachedGlobal narrative with selected country call-outs; statistic-level records keyed to topic sectionsGBD 2019 reference year; revised as later rounds landFraming sentences and citation checks, not modeling input
IHME GHDx GBD 2019 Smoking Tobacco Use Prevalence 1990-2019Age-standardized prevalence, smoker counts, percent change and cigarette-equivalents per capita - exposure beside burden, attributable deaths excluded by design204 countries and territories plus selected subnational units; country x sex x age group x yearAnnual 1990-2019, one fixed modeling roundThe exposure companion that pairs with a burden extract
Which measure answers which question
MeasureWhat it answersBest for
DeathsThe fatal toll - how many deaths smoking causes per geography, sex, age band and yearMortality attribution, cross-country league tables, headline journalism
Years of life lost (YLLs)How much earlier people die - mortality weighted by age at deathPremature-mortality costing, life-insurance exposure work
Years lived with disability (YLDs)The non-fatal toll - time lived in diminished healthMorbidity burden, disability-weighted planning
DALYsYLLs plus YLDs in one number - total burden on a single scaleHealth-economics cases, value-of-burden models, ESG screening
Population attributable fractionThe share of a disease's burden the risk factor explainsAttribution disputes, cause-specific decomposition
Summary exposure valueExposure strength on a comparable zero-to-one scaleExposure-outcome pipelines, forecasting covariates
The burden-plus-context join map
LayerRecordShared keysWhat it adds
Burden (outcome)IHME GBD Results Tool - Tobacco/Smoking Query InterfaceCountry or subnational unit, year, sex, age bandDeaths, DALYs, YLLs and YLDs with uncertainty bounds, 1990-2023
Exposure (modeled)World Bank - Prevalence of Current Tobacco Use (% of adults)Economy, yearAge-standardized any-tobacco prevalence, 2000-2024, with male and female splits
Exposure (surveyed)Eurostat - Daily Smokers of Cigarettes (sdg_03_30)Geography, sex, survey yearSelf-reported daily-smoking shares for 35 European entities, seven waves 2006-2023
PolicyWHO Report on the Global Tobacco Epidemic 2023Country, yearMPOWER achievement levels 1-4 per measure, back to 2008

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Tobacco Global - 204 countries and territories plus 660 subnational…

IHME GBD Results Tool - Tobacco/Smoking Query Interface

Tobacco Countries worldwide plus world and regional aggregates

Our World in Data - Smoking Dataset

Tobacco Global narrative with selected country call-outs (India

HealthData.org — Tobacco Disease & Injury Factsheets Data

Tobacco Global: 204 countries and territories plus selected…

IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019

location_name · sex · age_group_name …+4 more

Tobacco Roughly 265 economies, territories and aggregate regions…

World Bank - Prevalence of Current Tobacco Use (% of adults)

countryiso3code · date · value …+3 more

Tobacco All WHO member states (194+) plus regional and global aggregates

WHO Global Health Observatory Legacy Data Repository - Tobacco

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Are smoking-attributable deaths counted or modeled?

Modeled, like all attributable burden. The estimate contrasts observed mortality and disease burden with a counterfactual world where smoking disappeared, and publishes the difference as deaths, YLLs, YLDs and DALYs with 95% uncertainty intervals. The interval is part of the finding - a modeled figure quoted without its range is a claim, not a measurement.

Which dataset has GBD smoking attributable deaths by country?

The IHME GBD Results Tool - Tobacco/Smoking Query Interface carries them as queryable rows: deaths, YLLs, YLDs and DALYs for the Smoking risk factor across 204 countries and territories plus 660 subnational locations, every year from 1990 through 2023, each estimate with its uncertainty interval. Datadory delivers that cube shaped to your markets, beside Our World in Data's citation-ready chart series for the quotable headline.

Why do smoking-attributable death figures differ between sources?

Vintage and grain. The GBD 2023 release models every year through 2023 with full demographic detail; the HealthData.org topic record holds the GBD 2019 reference year for stable citation; harmonized chart series extend to 2022-2023 with point values only. Same underlying program in several cuts - so quote the vintage beside the number and never mix vintages inside one trend line.

Do the estimates carry uncertainty intervals?

Yes. Every row of the GBD cube ships lower and upper bounds of its 95% uncertainty interval beside the point estimate, and WHO's modeled prevalence observations carry their own bounds. Chart collections such as Our World in Data publish point values only, which is why analysts pair the two: the chart for the headline, the cube for the error bars.

Can smoking-attributable burden be joined to prevalence and policy data?

Cleanly, on shared country and year keys. The World Bank panel contributes age-standardized prevalence for roughly 265 economies from 2000 to 2024 with male and female splits, WHO's observatory adds modeled prevalence with bounds plus MPOWER policy grades, and Eurostat contributes surveyed European shares by sex. One rule holds the join together: compare the estimators, never splice them into a single series.