Datadory notebook

Smoking prevalence by country: one number per economy-year

Datadory delivers tobacco industry data covering smoking prevalence by country through the World Bank prevalence panel - indicator SH.PRV.SMOK, the share of adults aged 15+ currently using any tobacco product, age-standardized, for roughly 265 economies annually from 2000 to 2024 with male and female splits on identical keys - joined to Eurostat's European daily-smoker waves and WHO's ~300 modeled tobacco indicators - delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

Which dataset measures smoking prevalence by country?

World Bank - Prevalence of Current Tobacco Use (% of adults) — indicator SH.PRV.SMOK — is the number public-health arguments default to, and it earns the position structurally. One row per economy-year, roughly 265 economies, territories and aggregate regions, annual observations from 2000 to 2024 — about 6,625 rows — all carrying one definition: the percentage of people aged 15 and over who currently use any tobacco product, smoked or smokeless, daily or non-daily, age-standardized to the WHO Standard Population so a young median-age economy and an ageing one sit on the same scale.

Three properties make it unusually workable. Consistency: one indicator code applying one definition everywhere, which is precisely what cross-country tobacco work usually lacks. Reach: small states that commercial panels skip sit in the same table as the G20, keyed on ISO-3 codes. Completeness of context: cigarettes, pipes, cigars, cigarillos, waterpipes, bidis, kretek, heated tobacco and every form of oral or nasal smokeless tobacco count toward the figure; e-cigarettes stay out because they contain no tobacco. Within Datadory's catalog of 1,744 datasets across 159 viable industries, this record scores 10 out of 10 against a slice-wide mean of 7.81 — the only perfect score in the tobacco slice. Get a sample cut to your economies and years before anything else.

What does "smoking prevalence" actually measure in each record?

The phrase hides three different claims, and mixing them is the most common error in cross-country work.

  • The World Bank's SH.PRV.SMOK counts current use of any tobacco product among adults — smoked and smokeless, daily or non-daily, age-standardized. Broadest product set, widest geography.
  • Eurostat - Daily Smokers of Cigarettes (sdg_03_30) counts the population aged 15+ who smoke boxed cigarettes, cigars, cigarillos or a pipe — vapes and snuff deliberately excluded, so vaping cannot blur the combustible trend — measured only in Eurobarometer survey waves rather than interpolated annually.
  • WHO Global Health Observatory modeled estimates report current tobacco use with low/high uncertainty bounds attached to every value: Malaysia's 34.5% male estimate for 2022 ships as 34.5 [22.0-47.0], Cameroon's 9.1% for 2023 as 9.1 [6.5-11.8] — two bands that say more about underlying survey disagreement than any methodology note.

Pick one definition per analysis, document it, and never average across sources. A Eurostat share, a World Bank any-tobacco figure and a WHO modeled estimate can disagree about the same country in the same year without any of them being wrong — they are answering different questions.

What does one prevalence row look like?

One observation per economy-year, reproduced exactly as captured at research time:

# total-adults line
SH.PRV.SMOK     AFE  2023  11.8997926098489   # Africa Eastern & Southern
SH.PRV.SMOK     AFE  2022  12.1786708462139
SH.PRV.SMOK     AFE  2024  11.650004712267

# sex splits ride sibling codes on identical keys
SH.PRV.SMOK.MA  AFE  2023  19.9590625406671   # male adults
SH.PRV.SMOK.FE  AFE  2023   3.77511582966064   # female adults

Read the anatomy rather than the digits. Four coordinates fix every observation — indicator code, economy name, ISO-3 code, year — and the value carries full computational precision because each figure derives from harmonized survey estimates rather than rounded press releases. The gap between the male and female readings for the same region-year, nearly five to one here, is exactly the structural fact the sex-split variants exist to expose; the total line folds it away by construction. Economy-years nobody observed arrive null, and nulls ship explicit rather than imputed, so absence stays distinguishable from zero.

Do you need separate series for male and female smoking rates?

Yes — and build the split in from the start rather than deriving it later, because the sex gap is often the story itself.

The World Bank publishes dedicated indicators for each sex: Prevalence of Current Tobacco Use, Male (% of male adults) (SH.PRV.SMOK.MA) and the female counterpart SH.PRV.SMOK.FE, each with roughly 5,000 rows across the same 2000-2024 span as the 6,625-row total-population series, on identical economy-year keys — so a three-series panel assembles with a two-column join. The smaller row counts mean the disaggregated series carry more gaps; check balance before fitting a panel model.

Eurostat's sdg_03_30 splits every value into Total, Males and Females for 35 European entities aged 15+, roughly 620 observations across seven waves. In the 2014 wave the EU-27 read 27% overall — split 32% of men against 22% of women, a ten-point gap that is itself the headline for anyone forecasting combustible volume.

Which traps distort cross-country smoking comparisons?

Three failure modes account for most bad comparisons, and all three survive contact with good data.

Leaving aggregates inside a country ranking. The roughly 265 World Bank codes include regional composites such as Africa Eastern and Southern and income-group rollups riding the same schema as sovereign states. Any ranking that fails to tag them ranks Germany against an entire region. Every row carries its ISO3 code precisely so the filter stays one predicate long.

Averaging over unobserved economy-years. Null means not measured, never zero smoking. Panel completeness per year should be measured rather than assumed; average across unfiltered nulls and the missing states silently drag the panel toward nothing.

Mixing definitions mid-analysis. A Eurostat daily-cigarette share, a World Bank any-tobacco figure and a WHO modeled interval estimate are three different claims. Joined into one chart without labeling the units, they draw a series that looks continuous and is not.

How do the main smoking prevalence records compare?

Five records cover most country-level prevalence work. What separates them is geography, definition and grain rather than brand — the columns below are what decide which one anchors a project. As cataloged by Datadory, August 2026:

Who builds on country-level smoking prevalence?

Market researchers and consultants size nicotine demand from the total-adults line, then sharpen with the male and female splits — one definition doing the harmonization work that normally consumes a study's first month.

Investors and quant researchers read prevalence trajectories as the citable spine of tobacco-demand narratives across markets, on one definition rather than a patchwork of national statistics offices.

Data scientists and ML engineers drop a verified country-year panel with explicit nulls and stable keys into feature stores without a cleaning pass — the headline column arrives before the sex splits append onto the same keys.

Public-health and excise analysts set prevalence slopes beside policy scores and tax records, turning a league table into an evaluation of whether intervention moved behavior. WHO's Report on the Global Tobacco Epidemic 2023 contributes the MPOWER scores — every member state rated 1-to-4 on each demand-reduction measure back to 2008.

Every one of these jobs runs off the same typed rows — delivered daily, weekly, or hourly, your call — so the workflow is a sample request away rather than a pipeline build.

Why get smoking prevalence by country through Datadory?

Because the numbers were never the hard part — shape is. Prevalence here means three sibling indicator codes sharing one spine, aggregate regions hiding inside the country list, nulls that mean unmeasured rather than zero, and four definitions of 'smoking' spread across the field's catalogs. Reconciling all of that into one analysis-ready table eats days before any modeling starts.

Files, feeds, or straight into your warehouse. Daily, weekly, or hourly - your call. Datadory handles the decoding upstream of you: aggregates tagged so a sovereign-state filter stays one predicate, nulls kept explicit, the male and female splits aligned onto identical keys, and the Eurostat, WHO and IHME companions delivered joinable without a crosswalk. Every delivery ships with the field dictionary, sample rows for validation and a coverage statement stating exactly where the panel ends.

Start with a sample: name the economies, sexes and years you need, and real rows come back cut to them before anything recurring starts.

Where to go next

This page is one thread of the tobacco stack. Open the tobacco data hub for the complete scored index of the industry slice, or the ranked best tobacco datasets shortlist for the top of it. The cigarette consumption per capita page covers the demand side this prevalence panel pairs with naturally; the US cigarette production statistics page covers the American supply side. Dataset pages document the total-adults panel, the male and female variants and the WHO Global Health Observatory legacy repository down to the field dictionary, and the trade-off between the male-split and total lines is laid out in the male vs adults comparison. Practical order of operations: load the total-adults line for the cross-country baseline, append both sex splits on the shared keys, then layer MPOWER scores or GBD burden once the panel is clean.

Country-level smoking prevalence records compared (as of August 2026)
DatasetWhat it measuresCoverage and grainTemporal depthRole in the stack
World Bank - Prevalence of Current Tobacco Use (% of adults)Current use of any tobacco product, ages 15+, age-standardized; smoked and smokeless included, e-cigarettes excluded~265 economies plus regional and income aggregates; one row per economy-year; 6,625 rowsAnnual, 2000-2024The cross-country baseline everything benchmarks against
Eurostat - Daily Smokers of Cigarettes (sdg_03_30)Share of population aged 15+ smoking boxed cigarettes, cigars, cigarillos or a pipe, by sex; vapes and snuff excluded35 European entities incl. EU-27, EFTA and UK; ~620 observations, geography x sex x waveSurvey waves 2006, 2009, 2012, 2014, 2017, 2020, 2023European combustible-specific mirror
WHO Global Health Observatory Legacy Data Repository - TobaccoModeled current tobacco-use estimates with Low/High bounds, plus smokeless and e-cigarette series and MPOWER policy scores194+ member states plus regional aggregates; country x sex x year x indicator; ~300 tobacco indicatorsModeled estimates through the 2022-2023 roundsUncertainty-carrying breadth behind the headline panels
IHME GBD Results Tool - Tobacco/Smoking Query InterfaceModeled smoking prevalence and attributable burden - deaths, YLLs, YLDs, DALYs - with 95% uncertainty intervals204 countries and territories plus 660 subnational locations; location x age x sex x year x measureEvery year, 1990-2023Subnational depth and outcome linkage no survey panel offers

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Tobacco Roughly 265 economies, territories and aggregate regions…

World Bank - Prevalence of Current Tobacco Use (% of adults)

countryiso3code · date · value …+3 more

Tobacco Roughly 265 World Bank economies plus regional and income…

World Bank API - Tobacco Smoking Indicators (SH.PRV.SMOK*)

Tobacco ~265 countries, territories and aggregate regions on the World…

World Bank - Prevalence of Current Tobacco Use, Male (% of male adults)

Tobacco Roughly 265 countries, territories and aggregate regions…

World Bank - Prevalence of Current Tobacco Use, Female (% of female adults)

Tobacco 35 European geographies: EU-27 aggregate

Eurostat Daily Smokers of Cigarettes (sdg_03_30)

Tobacco All WHO member states (194+) plus regional and global aggregates

WHO Global Health Observatory Legacy Data Repository - Tobacco

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Which dataset measures smoking prevalence by country?

The World Bank prevalence panel - indicator SH.PRV.SMOK, Prevalence of Current Tobacco Use (% of adults). It holds about 6,625 rows covering roughly 265 economies annually from 2000 to 2024 on one age-standardized definition, with sibling indicators SH.PRV.SMOK.MA and SH.PRV.SMOK.FE carrying male and female splits of roughly 5,000 rows each on the same country-year keys.

Why do smoking prevalence figures differ between sources?

Because each answers a different question. The World Bank counts current use of any tobacco product among adults; Eurostat's sdg_03_30 counts people aged 15+ who smoke boxed cigarettes, cigars, cigarillos or a pipe; WHO GHO ships modeled estimates with uncertainty intervals; IHME models smoking tobacco use prevalence by country, sex and single age group. Pick one definition per analysis and never average across them.

How is country-level smoking prevalence delivered?

As typed rows - files, scheduled feeds, or straight into your warehouse, daily, weekly or hourly, your call. Every delivery carries the field dictionary, sample rows and a coverage statement alongside the values, so the schema you validate in the sample is the schema that ships into your models.

Do I need separate series for male and female smoking rates?

Yes, and they should arrive pre-split rather than derived later. The World Bank publishes dedicated indicators per sex - SH.PRV.SMOK.MA and SH.PRV.SMOK.FE, roughly 5,000 rows each over the same 2000-2024 span - while Eurostat disaggregates every value into total, males and females across its 35 European entities.

What traps distort cross-country smoking comparisons?

Three: mixing definitions that measure different products and populations, leaving World Bank regional and income aggregates inside a ranking of sovereign states, and averaging over unobserved economy-years as if a gap were a zero. Each is a filter away from fixed if the schema says which row is which.

How current is country-level smoking prevalence data?

It depends on the measure. The World Bank panel runs through the 2024 observation year; WHO's modeled estimates run through the 2022-2023 rounds; Eurostat's survey waves end at 2023 with seven readings since 2006. Any figure claiming this year's global prevalence is extrapolating, and should label itself accordingly.