Tobacco Data Provider: From the Prevalence Panel to the Excise Ledger · Head-to-head
GBD 2019 Smoking Prevalence Files vs Our World in Data Smoking
Which tobacco data provider: from the prevalence panel to the excise ledger data fits your job: IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019, or Our World in Data - Smoking. API, files, or your warehouse. Daily, weekly, or hourly.
IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019
Our World in Data - Smoking
Where the fields line up
No shared field names. These two answer different questions.
| Field | IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019 | Our World in Data - Smoking |
|---|---|---|
Location (country, territory or subnational unit) | documented | not in this set |
Sex | documented | not in this set |
GBD standard age group | documented | not in this set |
Reference year (1990-2019) | documented | not in this set |
Age-standardized smoking prevalence | documented | not in this set |
Number of smokers | documented | not in this set |
Percent change in age-standardized prevalence and smoker counts | documented | not in this set |
Cigarette-equivalents per capita | documented | not in this set |
95% uncertainty interval bounds | documented | not in this set |
entity | not in this set | Country or region the observation refers to. |
code | not in this set | ISO 3166-1 alpha-3 country code; a dedicated code (OWID_WRL) marks the world aggregate. |
year | not in this set | Reference year of the estimate. |
Coverage, side by side
| IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019 | Our World in Data - Smoking | |
|---|---|---|
| Geographic | 204 countries and territories plus selected subnational locations | All countries worldwide plus world and region aggregates |
| Temporal | Annual estimates, 1990-2019 | Varies by chart; roughly 1990 out to 2022-2023, prevalence running through 2022 |
| Granularity | Country x sex x age group x year | Country x year; some charts split by sex or age group |
What each contains
They tie on 1 attribute. Pick by fit, not by loyalty.
| IHME GHDx - GBD 2019 Smoking Tobacco Use Prevalence 1990-2019 | Our World in Data - Smoking | |
|---|---|---|
| Geography | 204 countries and territories plus selected subnational locations | All countries worldwide plus world and region aggregates |
| Temporal reach | Annual estimates, 1990-2019 | Varies by chart; roughly 1990 out to 2022-2023, prevalence running through 2022 |
| Granularity | Country x sex x age group x year | Country x year; some charts split by sex or age group |
| Measures | Age-standardized smoking prevalence, number of smokers, cigarette-equivalents per capita, percent change over the period | About 38 indicators: prevalence, daily smokers, cigarettes per smoker per day, sales per adult, death rates, secondhand smoke, affordability, taxes, bans, advertising enforcement, vaping in Britain |
| Scale and packaging | Six files totaling about 900 KB from one modeling system | Roughly 38 chart-backed tables, typically hundreds to thousands of rows each |
| Formats delivered | CSV estimate files plus a PDF information sheet | CSV and JSON table exports per chart |
| Best for | Citable age-sex estimates with uncertainty bounds inside a fixed 1990-2019 window | Dashboards mixing prevalence, deaths, prices and policy context in one place |
Or take both in one feed
They stack cleanly, and most serious tobacco work should. Build the long-run analytical spine on the GBD 2019 files - age-sex-country cells with intervals from 1990 to 2019 - then extend the headline trend to 2022 with the OWID prevalence series and layer its tax, ban, affordability and mortality charts around the result. Because Our World in Data harmonizes IHME GBD among its upstream systems, the two tell one continuous story rather than arguing with each other.
Two practical notes from the records. Mind the seam at 2019: the GBD side is a fixed modeling round whose values will not move, while the OWID tables restate whatever their upstreams last published, so anchor joins on Entity or Code plus year and expect small disagreements near the boundary rather than row-for-row identity. And remember the OWID tables carry point values only - once you cross back into the GBD files, every number regains its interval. Or take both in one feed.
API, files, or your warehouse. Daily, weekly, or hourly.
Fair questions
Is the GBD 2019 smoking prevalence dataset better than Our World in Data's smoking charts?
Better depends on the job. The GBD 2019 release wins on depth: 204 countries and territories split by sex and age group, with 95% uncertainty intervals, smoker counts and cigarette-equivalents per capita for 1990-2019. Our World in Data wins on breadth and recency: about 38 indicators running from prevalence to taxes and deaths, reaching 2022-2023.
Which of the two has more recent smoking numbers?
Our World in Data. Its prevalence series runs through 2022 and other charts reach 2023, while the GBD 2019 files stop at 2019 by design - they are the fixed record of one modeling round. For trends covering the 2020s the OWID tables are the only option of the two; for a defensible historical baseline the GBD files remain the citable source.
Do both datasets break smoking prevalence down by age group?
Not equally. Every GBD 2019 estimate carries a full sex-by-age-group structure using the standard GBD bands, so youth and adult cells compare directly. In the OWID collection the base shape is country by year, with only some charts splitting by sex or age group. If age cells are the requirement, the GBD files deliver them consistently.
Does either dataset include deaths caused by smoking?
Only Our World in Data's collection does: death rates from smoking and secondhand smoke, the share of deaths attributed to smoking, and lung cancer death counts sit among its charts. The GBD 2019 prevalence files hold prevalence, smoker counts and cigarette-equivalents per capita only - attributable deaths and DALYs live separately in IHME's GBD Results Tool.
Are the numbers in the two datasets consistent with each other?
Largely yes on prevalence, because Our World in Data harmonizes several of its smoking charts from IHME GBD material. Expect headline country shares to agree, with two cautions: the GBD side is frozen at its 2019 round while OWID restates later upstream revisions, and OWID publishes point values where GBD ships a 95% uncertainty interval around every estimate.
Which dataset should a market researcher sample first?
Both, in a specific order. Sample both, match each to its slide.