Datadory notebook
MEPS data: 493 files of US health spending and monthly coverage, delivered as rows
1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.
1,744 datasets. Pick your catch.
What is MEPS data?
The query sounds like a checkout decision; the answer is the closest thing American health economics has to a ground-truth instrument. Medical Expenditure Panel Survey (MEPS) is the Agency for Healthcare Research and Quality's continuously fielded survey of households, employers and medical providers, running every year since 1996, and it owns one question nobody else answers at this grain: what gets spent on health care in the United States, and who actually pays.
The published Household Component library spans 493 files across survey years 1996 through 2024: person, family, job, event and condition levels, plus full-year consolidated files, longitudinal panels, NHIS link files and projections aligned to the National Health Expenditure Accounts. It scores a perfect 10 on Datadory's quality rubric against a catalog-wide average of 7.81 across the 1,744 datasets we catalog - one of only two ceiling scores among the sixteen primary records in this slice. Datadory delivers the whole library as typed rows keyed on the person spine; get a sample cut to your years and service categories before anything else.
What do MEPS rows look like once delivered?
Person-grain rows after cleaning - one row per person, illustrative values on the real skeleton:
DUID PID DUPERSID PANEL YEAR AGE SEX RACETHX MONTHLY COVERAGE TOTEXP TOTSLF PERSON WT
---------------------------------------------------------------------------------------------------------
2320034 001 2320034001 27 2023 34 2 6 PRV x12 4,870 920 41,203
2320117 002 2320117002 27 2023 67 1 6 MCR x12 + PRV x12 31,450 2,140 12,876
2320149 001 2320149001 28 2023 8 1 1 MCD x12 1,610 0 58,442Which fields carry the weight?
Dictionaries are where MEPS projects quietly die, because the columns are coded, era-specific and load-bearing all at once. The spine is stable: DUID and PID identify the dwelling unit and the person within it, DUPERSID concatenates them into the join key every event file resolves back to, PANEL and DATAYEAR place the row in the rotating design, and AGELAST, SEX and RACETHX carry the edited demographics, with race/ethnicity harmonized to one comparable column across file years.
The signature variables are the coverage flags. Twelve per source per person-year - MCRmm23X for Medicare, MCAIDmmX for Medicaid/SCHIP, TRImm23X for TRICARE/CHAMPVA, PUBmm23X for any public plan, PEGmm23 for employer/union coverage - turning insurance from an annual label into a monthly timeline. Beside them sit the money pairs: TOTEXP and its splits by payer (self/family, Medicare, Medicaid, private, other) at event grain and as annual totals, with TOTSLF isolating out-of-pocket spending.
Then the trap worth naming out loud: the person weight changes name across eras - WTSAQ, WTPER and WTFINL appear in different file years - and any national estimate computed unweighted, or weighted with the wrong vintage, is simply wrong. Datadory ships each weight named, typed and documented per file year, so the lookup happens upstream of you rather than mid-analysis.
How wide does coverage run, and where does it stop?
Geography - the United States, civilian noninstitutionalized population: national estimates with census-region breakdowns. Nursing-home residents, the incarcerated and active-duty military sit outside the household universe by design, which is precisely when the Human Mortality Database or institution-side records earn their place in the stack.
Granularity - five grains (person, family, job, event, condition) plus consolidated, longitudinal and NHIS-link files, and twelve monthly coverage flags per person per year. Event files split by service category - dental, hospital inpatient, emergency room, office-based, outpatient, prescribed medicines, other medical expenses - so utilization studies should pick the family that matches the question rather than pulling blind.
Two scope notes belong in every plan. Employer-side microdata stay confidential and publish as summary tables at national, regional, state and metro levels, so firm-level premium questions answer at the aggregate the survey releases. And state-level detail on the household side lives in those same summary tables, not in the person files - a structural fact of the survey design that no amount of clever joining works around.
How is MEPS data delivered?
Name the years, service categories and grains when you request the sample and the extract arrives cut to that scope; the standing feed follows the same shape, so anything prototyped survives delivery intact. See how the slice behaves at machine speed in the life & health insurance data APIs, and how citation-grade work stays auditable in citation-grade research.
What can you build once MEPS is on site?
Four workflows pay for themselves fastest.
Coverage churn and continuity. Twelve monthly flags per person-year turn insurance history into transitions: who moved from employer coverage into Medicaid, how long uninsured spells ran, what continuity looks like by income band. An annual snapshot cannot see this structure at all, and most substitute feeds offer exactly that.
Pricing and utilization benchmarks. Spend distributions by event type and condition give health-plan actuaries an empirical backbone for plan-design and morbidity assumptions, and give investors covering managed care a national baseline to test a portfolio company's reported trends against.
Market sizing that survives diligence. Expenditure totals reconcile to the National Health Expenditure Accounts, so sizing work for payer-, provider- and pharmacy-adjacent markets carries a public denominator; the workflow continues in market-sizing.
Which datasets pair with MEPS?
MEPS prices care; four sibling records fill the edges it leaves open.
Stacked deliberately, they cover counts, context, mortality, carrier financials and the classroom - the standard configuration for this slice in best life & health insurance datasets.
Who builds on MEPS data?
Ranked by how directly the files answer the day job:
- Data scientists & modelers train expenditure and utilization models back through 1996 on survey-weighted targets no synthetic panel reproduces; the workflow ranks the slice in life & health insurance data for data scientists.
- Investors & quants covering managed care benchmark utilization frequency and payment mix against the national baseline before believing a portfolio company's trend line; see investors & quants.
- Market researchers & consultants size payer, provider and pharmacy-adjacent markets off denominators that reconcile to the National Health Expenditure Accounts; see market researchers.
- Journalists, academics & students cite national estimates whose methodology, response rates and variance guidance are published end to end; see journalists, academics & students.
- Actuarial and pricing teams at health plans read event-type spend distributions as the empirical backbone for plan-design pricing and morbidity assumptions.
Why get MEPS through Datadory?
Because the hard part was never the first extract - it is the tenth. Weight columns that rename themselves across eras. Enum definitions that shift between file years. Two panels hiding inside one calendar year. Event families that join to people only through one concatenated key. Each is survivable once; none is fun to re-solve in every new notebook.
Datadory normalizes before delivery: enums resolved against their file-year definitions, weights shipped named and typed per vintage, panel overlap flagged so deduplication is a column rather than a discovery, and event extracts pre-keyed to the person spine. Suppressed or lagging vintages surface explicitly rather than being found mid-model. When the next survey year lands, it arrives as new rows under the same dictionary - no re-integration project.
Name the years, service categories and grain when you request a sample and it arrives cut to that scope; the production feed follows the same shape, so anything prototyped on the sample survives delivery intact.
| Dimension | Coverage |
|---|---|
| Geography | United States, civilian noninstitutionalized population; national estimates with census-region breakdowns; employer-side tables extend to regional, state and metro levels |
| Temporal | Survey years 1996 through 2024 (29 annual waves); releases trail each survey year; overlapping two-year panels put two panels in most calendar years |
| Granularity | Person, family, job, event (per dental visit, admission, ER visit, office or outpatient encounter, prescription fill) and condition levels, plus consolidated, longitudinal and NHIS-link files; twelve monthly coverage flags per person-year |
| Methodology | Household reports reconciled against provider records through the Medical Provider Component; published codebooks, response rates and variance-estimation guidance behind every estimate |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
Medical Expenditure Panel Survey (MEPS)
Human Mortality Database (HMD)
U.S. Census Bureau Health Insurance Coverage Program
KFF State Health Facts and Health Policy Data
NAIC Insurance Data and Regulatory Tools
Kaggle Medical Cost Personal (Insurance) Dataset
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What comes out of a MEPS data delivery?
Typed rows rather than raw archives: consolidated full-year person files, event-level extracts by service category, twelve monthly coverage flags per person-year, expenditure totals split by payer, and sampling weights documented per file year. Name the years, categories and grain when you request a sample; it arrives in exactly that shape, and the standing feed follows it.
How far back does MEPS go, and how current is it?
Survey years 1996 through 2024 - twenty-nine annual waves, the longest continuous person-level series in the life & health insurance slice. Releases trail the survey year by design: the 2023 Full Year Consolidated file (HC-251) arrived in August 2025, 2024 survey-year files were scheduled for release between May and September 2026, and the 2025 employer-side tables complete the series July through November 2026.
How is MEPS data delivered?
API, files, or straight into your warehouse - daily, weekly, or hourly. Coded enums arrive resolved against their file-year definitions, weight columns ship named and typed per vintage, and event extracts come pre-keyed to the person spine on DUPERSID. The field dictionary is unchanged between the sample and the production feed, so validation takes minutes.