Tobacco · Institute for Health Metrics and Evaluation (IHME)

IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019

Datadory delivers ihme ghdx gbd 2019 smoking tobacco use prevalence 1990 2019 data covering modeled estimates of smoking tobacco use for 204 countries and territories from 1990 to 2019: age-standardized prevalence, estimated smoker counts, percent change across the window and cigarette-equivalents per capita, each cut by sex, age group and year with 95% uncertainty bounds attached to every point estimate, plus selected subnational locations, delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

What is the IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019 dataset?

Tobacco control's most-cited numbers, in their citable form. IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019 is the Global Health Data Exchange record hosting the estimate files behind the Global Burden of Disease Study 2019 smoking analysis published in The Lancet Public Health. Four estimate files carry the substance: age-standardized smoking prevalence, estimated number of smokers, percent change across the window, and cigarette-equivalents per capita. A codebook documents the columns and an information sheet states what the round does and does not claim.

Two construction details make this set behave differently from survey collections. First, the estimates are modeled - supply-side tobacco availability and consumption data fused with survey-based self-reported use - so thin national surveillance stops meaning holes in the panel. Second, every point estimate ships with the bounds of its 95% uncertainty interval attached, turning "how sure is this number" from a footnote into a column. What the files deliberately exclude is attributable disease burden: deaths and DALYs live in the companion GBD Results Tool collection, not here.

Within Datadory's catalog the record scores 9 out of 10 against a mean of 7.81 across 1,744 datasets. Get a sample of this dataset cut to the countries, sexes and age bands you work in.

What do sample rows look like?

One row per location-sex-age-year cell, and the spine never changes - only the meaning of the estimate column does, with the file it arrives in. Three cells in the delivered shape:

# Age-standardized smoking prevalence - share of the population smoking tobacco
location_name  : Brazil
sex            : Male
age_group_name : 30 to 34
year           : 2019
mean_estimate  : <point estimate>
lower_bound    : <95% uncertainty lower>
upper_bound    : <95% uncertainty upper>

# Number of smokers - people, not shares
location_name  : Japan
sex            : Female
age_group_name : 50 to 54
year           : 2000
mean_estimate  : <estimated smoker count>

# Cigarette-equivalents per capita - consumption proxy, per person
location_name  : Germany
sex            : Both
age_group_name : All ages
year           : 1990
mean_estimate  : <cigarette-equivalents per capita>

Read the anatomy rather than the placeholders. Dimensions identify the cell: country, sex, standard GBD age band, reference year. Measurement follows - a point estimate and, in the prevalence and counts files, the interval that travels with it. Because all four files share the spine, a thirty-year history for one sex-age slice of one country assembles with a filter, not a crosswalk. Live rows for the cells you name arrive with your sample, values exactly as the release publishes them.

What fields does the dataset include?

Six fields define the shared spine, each definition written against the release's codebook documentation during the August 2026 research pass and pinned against live records when your sample is cut. location_name identifies the country or territory, with selected subnational locations appearing beyond the 204 national units. sex sorts Male, Female or Both. age_group_name holds the standard GBD age band. year runs 1990 through 2019. mean_estimate carries the point estimate, and its meaning changes with the file: a share in the prevalence file, a headcount in the smokers file, a percent change in its own file, a per-person consumption figure in the cigarette-equivalents file. lower_bound and upper_bound bracket the 95% uncertainty interval around each point.

Two conventions worth knowing before you build. Treat mean_estimate as file-typed, not globally typed - a naive union of two files produces one column mixing shares and headcounts. And the uncertainty bounds are part of the datum, not decoration: inverse-variance weighting and interval propagation both start from those two columns. Anything deeper folds under additional fields on request below rather than being promised blind.

What does coverage look like across geography, time and granularity?

Geography - global by construction: 204 countries and territories, from the largest tobacco markets to small island states where national surveillance is thinnest and modeled estimates earn their keep. Selected subnational locations ride alongside the national units, which is unusual for a prevalence panel and worth asking about when scoping a sample.

Temporal - annual estimates from 1990 through 2019, produced as one fixed modeling round. That single-vintage character cuts both ways: methodology is internally consistent end to end, and there is a hard edge at 2019 - newer rounds exist in adjacent collections, not inside this one. The percent-change file summarizes movement across the whole window rather than adding years.

Granularity - country by sex by age group by year within the standard GBD age structure. Against the wider catalog this record scores 9/10, carried by complete global reach and an uncertainty treatment most prevalence panels never attempt.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel your stack already speaks. Rows arrive identical either way - cleaned, typed against the dictionary above, and carrying the vintage label so a 2019 estimate never masquerades as a current reading.

Every delivery ships the full field dictionary, validation rows keyed to the countries, sexes, age bands and years you named, and a schema that stays flat however many files you combine.

One scoping habit pays off: tell us which of the four estimate files matter - prevalence, smoker counts, percent change or cigarette-equivalents per capita - and the sample arrives shaped to exactly those columns.

Who uses this data, and for what?

  • Market sizing and demand forecasting - prevalence and smoker counts by sex and age give nicotine-demand models a thirty-year baseline; see our market researchers use cases page.
  • Actuarial and longevity work - smoking exposure resolved by cohort, sex and age band feeds mortality and morbidity assumption sets with quantified uncertainty.
  • Health-economics and simulation modeling - a consistent, uncertainty-tagged prevalence surface is the standard input layer for forecasting and cost models that predate burden attribution; see our data scientists use cases page.
  • Equity and portfolio research - prevalence trajectories are the volume-side backdrop behind tobacco-sector theses; see our investors quants use cases page.
  • Competitive intelligence - watch which markets' smoker populations are growing, aging or shrinking before positioning products in them; see our competitive intel use cases page.
  • Citation-grade sourcing - stories, theses and filings anchored to the published analysis itself rather than secondhand chart screenshots; see our journalists academics use cases page.

Which personas get the most value?

Data scientists and ML engineers (relevance 3/3) get a tidy global panel with uncertainty columns already attached - rare enough in epidemiology to be worth saying twice; see data scientists use cases. Market researchers and consultants (3/3) get sex- and age-resolved demand baselines for 204 countries in one consistent vintage; see market researchers use cases. Investors and quant researchers (2/3) get citable trajectory evidence behind volume-side arguments; see investors quants use cases. Competitive-intel and product teams (2/3) read market direction and slope before committing distribution; see competitive intel use cases. Journalists, academics and students (3/3) get the published analysis behind one of the field's most-cited figures, with the estimate files to check any claim against.

How does it compare to alternatives in its slice?

Within tobacco data, this record owns the citable epidemiology layer: modeled prevalence, smoker counts and per-capita consumption resolved by sex, age and year, uncertainty bounds attached, one consistent vintage end to end. The neighbors own different jobs.

IHME GBD Results Tool - Tobacco/Smoking Query Interface extends the same modeling program to attributable burden - deaths, DALYs, YLLs and YLDs - and reaches 2023; pair the two when the question moves from who smokes to what it costs. Our World in Data - Smoking curates roughly 38 charts spanning taxes, sales, affordability and vaping; breadth there versus depth here, and the head-to-head lays out the trade. World Bank API - Tobacco Smoking Indicators (SH.PRV.SMOK*) answers the simpler question - current tobacco use as a percent of adults, annually to 2024 - without age bands or uncertainty intervals. Where the question is a thirty-year, sex-and-age-resolved picture of who smokes and how much, this is the set that answers it.

What should I know before requesting a sample?

Four things worth settling upfront.

First, these are estimates, not censuses. Modeled values with stated uncertainty, not survey microdata or administrative counts - which is precisely what makes 204 countries comparable, and precisely why the bounds belong in your analysis rather than in a discarded column.

Second, the vintage has a hard edge. The window closes at the 2019 reference year and the round is fixed; anything needing post-2019 readings belongs to the GBD Results Tool collection, and we will say so at scoping rather than stretch this one past its design.

Third, scope by file, not just by geography. The four estimate files answer four different questions, and mean_estimate changes meaning between them - name the files and the columns arrive typed accordingly.

Fourth, attributable burden is a neighbor, not an inclusion. Deaths and DALYs sit in the GBD Results Tool collection on the same location-sex-age-year keys, so a combined extract is a join we handle at delivery. Say which cells matter - countries, sexes, age bands, years - and the sample comes back cut to exactly that shape.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - six shared-spine fields, one row per location, sex, age group and year
FieldTypeDefinitionExample
location_namestringCountry or territory the estimate refers to; selected subnational locations appear in addition to the 204 national units.Brazil
sexenumSex the estimate applies to: Male, Female or Both.Female
age_group_namestringAge band within the standard GBD age structure, from five-year groups such as 10 to 14 through the open-ended oldest band.30 to 34
yearintegerReference year of the estimate, running 1990 through 2019.2019
mean_estimatenumberPoint estimate of the file's measure - age-standardized prevalence, smoker count, percent change, or cigarette-equivalents per capita depending on which file the row ships in.see file-specific note
lower_bound / upper_boundnumberBounds of the 95% uncertainty interval accompanying each point estimate.interval around the point estimate

What teams do with it

  • Market sizing and demand forecasting Prevalence and smoker counts by sex and age give nicotine-demand models a thirty-year baseline instead of a single current snapshot.
  • Actuarial and longevity loading Smoking exposure resolved by cohort, sex and age band feeds mortality and morbidity assumption sets with quantified uncertainty rather than point guesswork.
  • Health-economics and simulation modeling A consistent, uncertainty-tagged prevalence surface is the standard input layer for cost-of-disease and forecasting simulations that predate burden attribution.
  • Cross-market benchmarking Rank 204 countries on age-standardized prevalence or per-capita cigarette-equivalents with intervals wide enough to keep the ranking honest.
  • Equity and portfolio research Prevalence trajectories are the volume-side backdrop behind tobacco equities and reduced-risk product theses - direction and slope, citably sourced.
  • Forecast model training Thirty annual observations per cell with lower and upper bounds give time-series and panel models both targets and uncertainty-aware loss weights.

Questions buyers ask

What does the IHME GHDx GBD 2019 smoking tobacco use prevalence data include?

Four estimate files from one modeling round: age-standardized smoking prevalence, estimated number of smokers, percent change in both measures across the window, and cigarette-equivalents per capita - plus a codebook documenting the columns and an information sheet stating what the round claims. Estimates span 204 countries and territories plus selected subnational locations, annually from 1990 to 2019.

How granular is the data?

Country by sex by age group by year. Age bands follow the standard GBD age structure - five-year bands such as 30 to 34 through the open-ended oldest group - and every estimate is available for males, females and both sexes combined. That is a far deeper cut than the country-year percentages most prevalence panels publish.

What are the 95% uncertainty bounds, and should I use them?

Every point estimate ships with lower and upper bounds of its 95% uncertainty interval - the modeling program's own statement of how sure it is. Treat them as part of the datum: weight by inverse variance, propagate intervals through simulations, or at minimum flag cells whose intervals are wide relative to the differences you are claiming.

Why do these figures differ from national survey numbers?

These are modeled estimates, built by combining supply-side tobacco availability and consumption data with survey-based self-reported use, adjusted to comparable definitions across countries and years. National surveys differ in question wording, age scopes and timing. When a figure disagrees with a local survey, the gap is usually definition rather than error - and the uncertainty bounds quantify the residual disagreement.

Does the data include smoking-attributable deaths or DALYs?

No - by design. This record carries prevalence and consumption measures only. Deaths, DALYs, YLLs and YLDs attributable to smoking come from the same modeling program through the GBD Results Tool collection, which reaches 2023; the two pair cleanly on the same location-sex-age-year keys.

Can a sample be cut to specific countries, sexes, age groups and years?

Yes. Name the countries, the sexes, the age bands and the year range - one market's full 1990-2019 history, a ten-country panel for adults, or a single sex-age slice across all 204 territories - and the sample arrives shaped to exactly that scope with the complete field dictionary attached. Samples precede any commitment.

Notes on this record

  • Scored 9/10 Datadory scores this record 9 out of 10 against a catalog mean of 7.81 across 1,744 datasets - complete global reach with uncertainty bounds on every point estimate.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing