World Bank - Prevalence of Current Tobacco Use, Female (% of female adults)

Datadory delivers Tobacco data covering world bank prevalence of current tobacco use female of female adults data - World Bank indicator SH.PRV.SMOK.FE: the share of women aged 15 and over who currently use any tobacco product, smoked or smokeless, daily or non-daily, age-standardized to the WHO Standard Population across roughly 265 countries, territories and aggregate regions for every year from 2000 through 2024, about 5,000 country-year observations - delivered daily, weekly, or hourly.

What is the World Bank prevalence of current tobacco use, female (% of female adults) dataset?

It is World Bank indicator SH.PRV.SMOK.FE from the World Development Indicators: the percentage of the female population aged 15 and over who currently use any tobacco product, on a daily or non-daily basis. The definition is deliberately wide on products and narrow on substance. Cigarettes, pipes, cigars, cigarillos, waterpipes (hookah, shisha), bidis, kretek, heated tobacco products and all forms of smokeless tobacco, oral or nasal, all count. E-cigarettes do not - they contain no tobacco, so vaping cannot inflate a series whose job is to size the actual tobacco-consuming base.

Two properties make the numbers travel. Each value is age-standardized to the WHO Standard Population, so a country with young adult women and a country with older ones sit on identical footing before you compare them. And the underlying compilation comes from the WHO Global Health Observatory, published through World Bank Data alongside the rest of the WDI development canon - which means this series joins cleanly onto thousands of other country-year indicators already keyed to the same economy list.

For the Tobacco industry that adds up to roughly 265 countries, territories and aggregate regions, every year from 2000 through 2024, about 5,000 country-year observations in this gender-split variant. Get a sample of this dataset and it comes back cut to the markets and years you name.

What does a sample of this dataset look like?

Rows exactly as they arrive:

# one observation = country x year, females aged 15+
indicator.id  : SH.PRV.SMOK.FE     # prevalence of current tobacco use,
                                   # female (% of female adults)

country.value    : Africa Eastern and Southern
countryiso3code  : AFE             # ISO-3166 alpha-3 economy code
date             : 2023            # observation year
value            : 3.775           # 3.78% of adult women currently use
                                   # any tobacco product

country.value    : Africa Western and Central
countryiso3code  : AFW
date             : 2023
value            : 1.141           # 1.14% of adult women

Two rows, three habits worth keeping. First, the identifier travels with every row - indicator code, economy name, ISO-3 alpha-3 code, year - so the panel self-assembles with a group-by and never needs fuzzy country-name matching. Second, notice what these particular rows are: aggregate regions, not countries. The ~265-entry economy list mixes sovereign states with regional groupings, so decide up front whether Africa Eastern and Southern belongs beside Kenya in your panel or gets aggregated out of it. Third, the values carry full precision - 3.77511582966064 arrives as 3.77511582966064 - with a recommended display rounding carried separately, so no information was destroyed upstream to make the column look tidy.

Your sample returns this exact shape scoped to the economies and years you name; the schema in the sample is the schema that ships.

Which fields does the field dictionary define?

Eight fields carry the entire series, every one of them verified against published column definitions. Five of them do the analytical work - country.value and countryiso3code identify the economy (full name plus stable alpha-3 code, so joins survive renames), date holds the observation year, value carries the age-standardized percentage itself, and indicator.id pins every row to SH.PRV.SMOK.FE so a file that quietly gained the male or total series announces itself instead of poisoning your panel.

The remaining three are constants and flags worth understanding rather than ignoring: decimal gives the recommended display rounding (1 for this indicator), unit ships empty because the percentage is already baked into the indicator definition, and obs_status carries observation-status markers where they exist. Constants are a design decision, not missing metadata - they guarantee any two values you compare were produced on identical terms.

Additional fields on request: the published grid stops at these eight columns. Derived conveniences - income-group or region labels, aggregate-versus-sovereign flags, the male and total-population series aligned into one sex-split panel, null-gap summaries per economy - get enumerated, defined and named for your pipeline when your sample is cut. Say what your model needs at scoping; column naming locks there.

Where does coverage run, and at what grain?

  • Geo: roughly 265 entities on one economy list - sovereign states, dependent territories and aggregate regions such as Africa Eastern and Southern. One code scheme for all of them means a global pull needs no per-region mapping work, and filtering aggregates back out is a single predicate on the alpha-3 code set.
  • Temporal: every year from 2000 through 2024 - twenty-five annual observations per economy, deep enough to carry trend features across two and a half decades of tobacco-control change. Absent readings stay absent rather than being interpolated to force continuity.
  • Granularity: one observation per country x year, women aged 15+, expressed as a percentage of female adults - about 5,000 rows in this gender-split variant. Small enough to email, structured enough to join onto anything with an economy column.

How is the data delivered through Datadory?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel your stack already speaks and set the cadence to match the decision being fed. Five thousand rows is a payload that fits anywhere - a dashboard widget pulling hourly, overnight files loading into BI, or a direct landing in Snowflake, BigQuery or Redshift beside your sales and excise tables. Cadence changes are a settings conversation, not a re-integration project.

Every delivery ships with the field dictionary above unchanged and validation rows included, one row per observation. Name the economies and years when you request the sample; it comes back cut to that exact shape before any commitment.

Who builds on this dataset?

Ranked by how directly one row settles the day job:

  1. Tobacco strategists and category planners. A female consumer base sized economy-by-economy on one stable definition - the demand-side denominator underneath volume forecasts, and the only version of the number that counts smokeless users in markets where women chew rather than smoke.
  2. Investors and quant researchers covering consumer staples. Female prevalence as a slow-moving regulatory-and-demand input: twenty-five annual observations is deep enough for trend features and clean enough to regress against tobacco-sector revenue by market.
  3. Public-health and NCD program teams. Sex-disaggregated prevalence against which cessation and prevention programs targeting women get planned and audited - the same age-standardized arithmetic the global health institutions reason over.
  4. Insurers and actuaries. Country-level female tobacco-use rates feeding mortality and morbidity assumptions for life and health books, at portfolio grain rather than world-average hand-waving.
  5. Data scientists joining WDI panels. The mined pattern is the simplest one: SH.PRV.SMOK.FE keys straight onto any country-year WDI extract as a demographic feature, no harmonization layer required.

Set against the wider Datadory catalog - average quality score 7.81 across 1,744 datasets - this slice scores 10/10, carried by verified field documentation and a definition held stable across the full quarter-century.

Which personas get the most value?

Investors and quants get a sex-disaggregated demand input with documented semantics and full-precision values, ready for factor construction without cleanup. Data scientists get enum-clean identifiers - alpha-3 codes, pinned indicator IDs, typed numerics - so the panel loads in pandas in minutes and joins onto any other World-Bank-keyed table without fuzzy matching. Market researchers and consultants get citable, age-standardized percentages for sizing decks, defensible because the standardization method is stated. Journalists and academics get the version of the female prevalence numbers that international institutions publish and defend. Across all four the constant holds: one narrow definition, measured identically, twenty-five years deep.

What should you know before requesting a sample?

Three things, stated up front. First, these are aggregate percentages, not microdata - no individual respondent records exist here, so person-level analysis needs a different instrument. Second, the economy list mixes sovereign states with aggregate regions - tell us which you want and we cut the sample accordingly, because a panel that silently contains both double-counts. Third, nulls mean not observed: choose whether you want the gaps kept honest or bridged against companion series, because those two choices produce different panels and both are legitimate.

Start at Get a sample of this dataset.

Notes that pair well with this page:

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - the eight fields carrying SH.PRV.SMOK.FE
fieldtypedefinitionexample
indicator.idstringIndicator code; pinned to SH.PRV.SMOK.FE on every row so a mixed-series file announces itself.SH.PRV.SMOK.FE
country.valuestringCountry, territory or aggregate region name, on the World Bank economy list.Africa Eastern and Southern
countryiso3codestringISO-3166 alpha-3 economy code; the stable join key.AFE
datestringObservation year.2023
valuenumberAge-standardized prevalence of current tobacco use among females aged 15+ (% of female adults); null when not observed.3.77511582966064
decimalintegerNumber of decimal places recommended for display; the stored value keeps full precision regardless.1
unitstringUnit of measure; ships empty because the percentage is fixed in the indicator definition.
obs_statusstringObservation-status flag marking estimate or missing conditions where present.

Coverage - geography, temporal range, granularity

DimensionCoverage
GeographyRoughly 265 countries, territories and aggregate regions on the World Bank economy list - sovereign states, dependencies and regional groupings under one alpha-3 code set
TemporalEvery calendar year from 2000 through 2024 - twenty-five annual observations per economy
GranularityOne observation per country x year, females aged 15+, percentage of female adults; about 5,000 rows for this gender-split variant

Questions buyers ask

What exactly does SH.PRV.SMOK.FE measure?

The percentage of the female population ages 15 years and over who currently use any tobacco product on a daily or non-daily basis. Current use spans smoked tobacco (cigarettes, pipes, cigars, cigarillos, waterpipes, bidis, kretek), heated tobacco products and every form of smokeless tobacco, oral or nasal. E-cigarettes are excluded because they contain no tobacco.

Which years does the female prevalence series cover?

Every calendar year from 2000 through 2024, one observation per country and year - roughly 5,000 country-year rows for this gender-split variant across approximately 265 countries, territories and aggregate regions. Values are absent rather than zero where no underlying measurement exists.

Why do some country-years return null instead of a number?

Null means not observed. The series carries the absence honestly instead of interpolating or filling with zeros, so a missing reading for a small economy never masquerades as proof that no woman there uses tobacco. Filter or bridge the nulls deliberately depending on whether your model wants continuity or fidelity.

Does the figure include chewing tobacco and vapes?

Chewing and all other smokeless tobacco are inside the definition; vapes are outside it. Because e-cigarettes contain no tobacco they do not count, so the series isolates the actual tobacco-consuming base regardless of delivery form - which matters most in markets where women chew rather than smoke.

How is the female series different from the male and total variants?

Same definition, same economy list, same 2000-2024 timeline - only the population differs: women here, men in the male variant, everyone in the total-population series. Requested together they form a sex-split panel on identical terms, so the male-female gap becomes a computed column rather than a reconciliation project.

Is cross-country comparison legitimate given different age structures?

Yes - each value is age-standardized to the WHO Standard Population before publication, so a country with young adult women and one with older women sit on identical footing. Differences in the published percentages reflect behavior rather than demography.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing