World Bank - Prevalence of Current Tobacco Use (% of adults)

Datadory delivers tobacco data covering World Bank indicator SH.PRV.SMOK, prevalence of current tobacco use among adults: one age-standardized percentage per economy per year for roughly 265 countries and aggregates from 2000 to 2024, about 6,625 observations on a single WHO-harmonized definition, delivered daily, weekly, or hourly.

What is the World Bank prevalence of current tobacco use (% of adults) dataset?

The number public-health arguments default to. Indicator SH.PRV.SMOK in the World Development Indicators reports the percentage of the population ages 15 and over who currently use any tobacco product - smoked or smokeless, daily or non-daily - age-standardized to the WHO Standard Population so a young median-age economy and an ageing one sit on the same scale. The count includes cigarettes, pipes, cigars, cigarillos, waterpipes (hookah, shisha), bidis, kretek, heated tobacco products and every form of oral or nasal smokeless tobacco; e-cigarettes are excluded because they contain no tobacco. The series reaches the World Bank catalog from the WHO Global Health Observatory, one compiler applying one definition across all reporters.

Three properties make the panel unusually workable. Scale: roughly 265 economies, territories and aggregate regions. Reach: annual observations from 2000 to 2024, about 6,625 rows for this total-adults variant. Consistency: one indicator code carrying one definition everywhere, which is precisely the property cross-country tobacco work usually lacks. Within Datadory's catalog of 1,744 datasets across 159 viable industries, this record scores 10/10, carried by fully verified field documentation and near-global coverage. Get a sample of this dataset cut to your economies and years before anything else.

What do sample rows look like?

One row per economy-year, reproduced exactly as captured at research time:

# one observation - economy x year
indicator.id     : SH.PRV.SMOK   (Prevalence of current tobacco use (% of adults))
country.value    : Africa Eastern and Southern
countryiso3code  : AFE
date             : 2023
value            : 11.8997926098489      # percent of adults, age-standardized

# the same spine, adjacent vintages
AFE  2022  12.1786708462139
AFE  2024  11.650004712267

# sex splits ride sibling indicator codes on identical keys
SH.PRV.SMOK.MA   AFE 2023  19.9590625406671    # male adults
SH.PRV.SMOK.FE   AFE 2023   3.77511582966064   # female adults

Read the anatomy rather than the digits. Four coordinates fix every observation - indicator code, economy name, ISO-3 code, year - and the value carries full computational precision because each figure derives from harmonized survey estimates rather than rounded press releases. The gap between the male and female readings for the same region-year, nearly five to one here, is exactly the structural fact the disaggregated variants exist to expose; the total line folds it away by construction. Economy-years that went unobserved arrive null, and nulls ship explicit rather than imputed, so absence stays distinguishable from zero.

What fields does the dataset include?

Eight documented fields, two jobs. The identification band (indicator.id, country.value, countryiso3code, date) fixes every observation in place; the payload band carries value plus the qualifiers that make it trustworthy - unit (empty for this indicator, since percent of adults is fixed by the definition itself), obs_status for estimate and missing-value markers, and decimal for display precision.

Additional fields on request. Everything outside the verified core is where scoping happens: the companion series sharing this spine - female (SH.PRV.SMOK.FE) and male (SH.PRV.SMOK.MA) prevalence on identical economy-year keys - the wide layout with one row per economy and one column per year, and pre-cut regional and income-group aggregates. Those get confirmed against live records when your sample is cut rather than promised blind.

Where does coverage reach?

  • Geo: roughly 265 economies, territories and aggregate regions on the World Bank country list - individual countries from the largest to the smallest sitting beside pre-computed regional and income-group rollups, every one keyed on ISO-3 codes.
  • Temporal: annual observations from 2000 through 2024. Per-economy windows vary with national survey practice, so aligning start years is part of any serious comparison - and the dictionary makes that alignment mechanical rather than archaeological.
  • Granularity: one observation per economy per year for the total-adults measure; the male and female splits hold the same grain under their own indicator codes, so a three-series panel assembles with a two-column join.

That footprint is why this record anchors the tobacco data hub. Set against the wider catalog - an average quality score of 7.81 across all 1,744 datasets - this slice scores 10/10, held up by verified definitions and coverage breadth rather than by novelty: it is the measured baseline other tobacco series get benchmarked against.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel your team already works in and set the cadence to match the decision being fed: a flat country-year panel refreshed on schedule for quarterly strategy reviews, a direct load into Snowflake, BigQuery or Redshift for teams running prevalence as model input, or point lookups for research workflows. Because every delivery is one economy-year grid keyed on ISO codes, it joins cleanly against the rest of the catalog - tobacco excise tax records for the policy layer, Global Burden of Disease estimates for outcomes.

Every delivery ships with the field dictionary above unchanged, sample rows for validation, and nulls left explicit rather than imputed. Sample first: name the economies, the years and the layout - long or wide - and the extract arrives shaped to them, with the sex splits attached where the analysis needs them, before any commitment.

Who builds on it?

Ranked by how directly one economy-year percentage settles the day job:

  1. Market researchers and consultants. Market sizing starts from the total-adults line, then sharpens with the male and female splits - one definition doing the harmonization work that normally consumes a study's first month. Patterns sit on the market researchers use cases page.
  2. Data scientists and ML engineers. A verified country-year panel with explicit nulls and stable keys drops into feature stores without a cleaning pass; the headline column arrives before the sex splits append onto the same keys. Workflows live on the data scientists use cases page.
  3. Investors and quant researchers. Prevalence trajectories give tobacco-demand narratives a citable spine across markets, on one definition rather than a patchwork of national statistics. Background on the investors quants use cases page.
  4. Journalists, academics and students. The standard adult-prevalence measure, documented to type and example, turns "smoking is falling" into a citable series with named provenance. Citation patterns sit on the journalists academics use cases page.
  5. Developers and data-product builders. One indicator code scoped by country and year range powers quick widgets and dashboards without custom parsers. Integration angles are on the developers builders use cases page.

Which notes and neighboring datasets pair with it?

Provenance note - compiled into the World Development Indicators from the WHO Global Health Observatory, the same program behind WHO's wider tobacco-control collections. One upstream estimator feeds both the total line and the sex splits, which is why the three never disagree about the same economy-year.

Completeness note - Datadory scores this record 10/10 with all eight fields documented to type and example. What sits outside the verified core is listed above under additional fields on request rather than approximated here.

Where to go next - inside Tobacco, the female variant and the male variant hold the same panel split by sex, and the indicator-family record covers all three codes together. The trade-off between the male-split and total lines is laid out in the male vs adults comparison. For outcome-side depth, the IHME GBD Results Tool adds attributable burden, the WHO Global Health Observatory legacy repository widens to some 300 tobacco indicators including MPOWER policy scores, and Our World in Data's Smoking topic packages charts with provenance notes. Start from the best tobacco datasets ranking or the tobacco data guide to see where this record lands in the slice.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary for World Bank - Prevalence of Current Tobacco Use (% of adults) - eight verified fields
FieldTypeDefinitionExample
indicator.idstringIndicator code identifying the series on every record.SH.PRV.SMOK
country.valuestringCountry, territory or aggregate region the observation belongs to.Africa Eastern and Southern
countryiso3codestringISO-3166 alpha-3 code - the join key most pipelines land on.AFE
datestringObservation year; one record per economy-year.2023
valuenumberPrevalence of current tobacco use, percent of adults, age-standardized; null where the economy-year is unobserved.11.8997926098489
unitstringUnit of measure qualifier; empty for this indicator because percent of adults is fixed by the definition.
obs_statusstringObservation status flag separating estimated or provisional values from finals.
decimalintegerRecommended number of decimal places for displaying the value.1

Sample rows - Africa Eastern and Southern aggregate as captured at research time

country.valuecountryiso3codedatevalue
Africa Eastern and SouthernAFE202212.1786708462139
Africa Eastern and SouthernAFE202311.8997926098489
Africa Eastern and SouthernAFE202411.650004712267

Coverage chips - geography, temporal depth and granularity

DimensionCoverage
Geography~265 economies, territories and aggregate regions, ISO-3 coded, with pre-computed regional and income-group rollups
TemporalAnnual observations 2000-2024; per-economy windows vary with national survey practice
GranularityOne observation per economy per year; male and female splits hold the same grain on sibling indicator codes

Questions buyers ask

What does the SH.PRV.SMOK indicator measure exactly?

The percentage of the population ages 15 and over who currently use any tobacco product - smoked or smokeless, on a daily or non-daily basis - age-standardized to the WHO Standard Population. Cigarettes, waterpipes, bidis, kretek, heated tobacco and all oral or nasal smokeless products count; e-cigarettes do not, because they contain no tobacco.

How many countries does the panel cover, and how deep does it run?

Roughly 265 economies, territories and aggregate regions, observed annually from 2000 through 2024 - about 6,625 rows for the total-adults variant. Per-economy windows vary with national survey practice, so a sample cut to named countries comes back with each window mapped explicitly before any modeling starts.

Can male and female prevalence be pulled separately?

Yes. Sibling indicator codes cover female adults and male adults on identical economy-year keys, about 5,000 observations each, so a three-series panel assembles with a simple join. Samples ship the total line and both splits side by side on request, which is the fastest way to test whether disaggregation changes your story.

Why are e-cigarette users not counted in the figures?

Because e-cigarettes and similar devices contain no tobacco, they fall outside the indicator's definition even where usage is widespread. Heated tobacco products do count. For markets where vaping substitutes heavily for smoking, treat the series as a tobacco-use measure rather than a nicotine-use measure and say so explicitly in methodology notes.

What does a null value mean in this panel?

An economy-year that went unobserved - not a zero prevalence rate. Nulls ship explicit rather than imputed, and the obs_status qualifier separately marks estimated figures, so the difference between absent and approximate survives into your pipeline instead of silently flattening during a load.

What should be settled before requesting a sample?

Four things. Which economies matter, so per-economy window depth comes back mapped. Whether the male and female splits belong beside the total line. Long or wide layout for the warehouse landing. And the join key your systems expect - ISO-3 codes ride on every row, which covers most pipelines as-is.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing