SHED Public Use Data Files (2013-2025)
Datadory delivers shed public use data files 2013 2025 data - the complete Federal Reserve SHED microdata archive: thirteen annual waves plus two 2020 COVID supplements, fifteen packages in all, 815 variables on 12,934 adult respondents in the 2025 file, a four-weight estimation system, and per-wave codebooks documenting every retained column.
What are the SHED Public Use Data Files (2013-2025)?
The SHED Public Use Data Files (2013-2025) are the raw respondent-level records behind the Federal Reserve Board's annual reading of how American households are doing financially: fifteen wave packages - thirteen annual surveys from 2013 through 2025 plus April and July 2020 supplements fielded on the 2019 panel - each shipping its questionnaire responses, weight system and codebook as one self-contained study.
Most teams meet SHED through a headline: some share of adults could not cover a $400 emergency expense. The published report averages that away. The files keep it split. In the 2025 wave the emergency-savings item (EF1) lands 5,607 no against 7,327 yes, and each of those 12,934 respondents carries the complete demographic and financial profile behind their answer - age, income band, education, state, banking status, housing tenure, gig work. One row per adult, 815 variables in the current file, every value labeled in that wave's codebook.
Get a sample of this dataset and see real SHED rows with their wave documentation before you commit.
What do SHED records look like?
Three respondents pulled from recent waves, unchanged from the released files: an Arizonan, 84, on a middle income who would charge a $400 surprise and pay the card in full at the next statement; a New Jerseyan earning $150,000 or more who would not; a Washingtonian without a high school diploma who holds a bank account anyway.
# SHED public use file - selected variables, one row per adult respondent
| shedid | ppage | ppinc7 | ppeduc5 | ppstaten | EF1 | EF3_a | BK1 | weight |
| --------- | ----- | ------------------ | ---------------------------------- | -------- | --- | ----- | --- | ------ |
| 202304484 | 84 | $50,000 to $74,999 | Some college or Associate's degree | az | Yes | Yes | Yes | 0.6467 |
| 202204577 | 80 | $150,000 or more | Master's degree or higher | nj | Yes | No | Yes | 0.9687 |
| 202301342 | 60 | $75,000 to $99,999 | No high school diploma or GED | wa | Yes | No | Yes | 1.0129 |Three details repay attention. EF1 and EF3_a are not abstract indices - they carry the exact question wordings behind the most quoted statistic in household-financial-resilience research, captured per respondent rather than averaged into a press release. ppstaten enters the files from the 2018 wave onward, which is where most commercial value hides: regional cuts of any well-being measure become possible the moment state arrives on the row. And weight is what turns 12,934 sampled adults into a nationally representative picture - ignore it and every percentage you compute is quietly wrong.
What fields does each SHED wave carry?
Twelve documented entries below span both layers of the archive: the wave packages themselves and the variables inside them. The 2025 file alone runs to 815 columns, grouped into banking (BK*), credit access and behavior, the emergency-expense series (EF*), income and expenses, employment and gig work, care work, education and student loans, housing (GH*, R*) and retirement, with derived demographic flags closing each record.
Every retained variable carries its question wording, type, range, missing count and labeled response values in that wave's codebook, so nothing in the table below requires reverse engineering:
Where does SHED coverage run, and at what grain?
- Geography: United States, nationally representative of adults, with respondent state identifiers in every wave from 2018 onward; finer geography is withheld from the files by design, so metro-level cuts need a companion dataset.
- Temporal: fifteen packages - annual waves 2013 through 2025, plus the April 2020 and July 2020 supplements fielded on the 2019 panel specifically to catch pandemic-period disruption while it was happening.
- Granularity: one record per adult respondent - 12,934 records by 815 variables in the 2025 wave - with a four-weight system (
weight,weight_pop,panel_weight,panel_weight_pop) supporting both single-year population estimates and cross-year longitudinal work.
Set against the wider Datadory catalog - average quality score 7.81 across all 1,744 datasets - this slice scores 9/10, carried by complete per-wave documentation and a thirteen-year unbroken series that no private panel matches at this depth.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
You pick the channel and the cadence; the wave structure above travels unchanged across all three. Full-archive pulls suit research teams rebuilding thirteen waves of trend lines once and joining them to proprietary panels. Wave-, state- or module-scoped extracts suit product surfaces that render well-being benchmarks inside an app. Warehouse delivery suits analysts keeping weighted SHED measures beside their own transactional tables in SQL. Cadence changes are a settings conversation, not a re-integration project, and every sample arrives mapped to the exact dictionary above.
Who uses this data, and for what?
- Trend lines with receipts - reconstruct any well-being measure across thirteen waves and say whether conditions improved or eroded, computed from respondent records rather than quoted from successive press releases.
- Segmentation and sizing - weight the $400-expense and three-month-savings answers by income band, age, state and education to size demand for liquidity products instead of buying a custom incidence study.
- Replication and verification - reproduce any figure in the Fed's annual reports from the underlying records, the standard first step when a client deck or regulator asks where a number came from.
- Longitudinal inference - exploit prior-year case identifiers in the 2018 and 2019 files and the panel weights to follow the re-interviewed subsample across years rather than treating each wave as a fresh draw.
- Model training and NLP - hundreds of cleanly labeled categorical variables paired with outcome flags like banking status make a rare supervised-learning substrate, and a benchmark for scoring synthetic survey responses.
- Policy research and journalism - quantify unbanked rates, hardship exposure and gig-work prevalence with citable respondent-level evidence behind every central-bank headline.
Get a sample of this dataset scoped to your waves, states and question modules.
Which personas get the most value?
Market Researchers & Consultants get the definitive household-balance-sheet segmentation base: when a deck claims X percent of adults cannot absorb a $400 surprise, this is the file that reproduces the claim, sliced however the client sells. Competitive Intelligence & Product Teams get demand-side evidence for savings, credit and banking products straight from consumers rather than from their own funnels - including where account ownership breaks down by state and income. Data Scientists & ML Engineers get a documented, codebook-labeled microdata archive with a stable weight system, ready to join onto panels or feed propensity models. Developers & Data-Product Builders get one consistent wave-by-wave schema to integrate once and reuse across every year thereafter. Journalists, Academics & Students get the citable ground truth under every Federal Reserve well-being statistic, with the documentation to defend each disaggregation.
Which notes pair with this dataset?
Notes and adjacent reading:
- Consumer Finance data hub - the pooled view of the industry, from complaint archives to household surveys.
- Best consumer finance datasets - where this archive sits in the ranked shortlist.
- Federal Reserve Survey of Household Economics and Decisionmaking (SHED) - the analytical twin: same thirteen waves, organized around the variable dictionary and published findings rather than the file-by-file record.
- vs USDA Food Access Research Atlas - a worked comparison of person-level survey microdata against tract-level place data.
- FDIC National Survey of Unbanked and Underbanked Households - the biennial counterpart with overlapping banking-status measures, useful for cross-validating any unbanked estimate across two federal instruments.
- Board of Governors of the Federal Reserve System source profile - the institution behind SHED, and everything else it releases.
Browse the whole vertical on the Consumer Finance data hub or the ranked shortlist of the best consumer finance datasets.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
Wave package | text | Each year ships as one self-contained study: questionnaire responses, a four-variable weight system and a PDF codebook. Fifteen packages cover 2013-2025 plus the two 2020 supplements. | 2025 wave |
public2025.csv columns | text | The 2025 wave's respondent table: 815 columns spanning banking (BK*), credit access and behavior, the EF emergency-expense series, income, employment and gig work, care work, education and student loans, housing, retirement, and derived demographic flags. | 815 |
shedid | string | Unique respondent identifier introduced with the 2025 file; earlier waves use wave-specific case IDs, with prior-year case IDs in the 2018 and 2019 files enabling cross-year links. | 202304484 |
weight / weight_pop | number | Respondent analysis weight and population-scaled post-stratification weight shipped in every recent wave; weighted shares reproduce the published national figures. | 0.6467 / 13225.34 |
panel_weight / panel_weight_pop | number | Panel counterparts of the two main weights, covering the longitudinal subsample re-interviewed across consecutive years. | <returned in your sample> |
ppstaten | string | Respondent state of residence, included in the public files from the 2018 wave onward - fifty states plus DC. | az |
prior-year case IDs | string | Linking identifiers added to the 2018 and 2019 files that let analysts follow the same respondent across years. | 2018 case id |
race_5cat / ppracem | enum | Race variables carrying a dedicated Asian category, appended retrospectively to the 2013-2019 files and corrected for missing values in mid-2026. | Asian |
EF1 | boolean | Whether the respondent has set aside emergency or rainy-day funds covering three months of expenses; the 2025 codebook tabulates 5,607 no against 7,327 yes. | Yes |
EF3_a | boolean | Whether the respondent would put a $400 emergency expense on a credit card and pay the statement in full - the item behind the best-known resilience benchmark. | No |
BK1 | boolean | Whether the respondent and/or spouse or partner holds a checking, savings or money market account - the basis for banked and unbanked classification. | Yes |
codebook PDFs | text | Per-wave documentation giving every retained variable's question wording, type, range, missing counts and labeled response values. | EF1: No=5607, Yes=7327 |
Questions buyers ask
How many SHED waves exist, and what ships with each one?
Fifteen packages: annual waves from 2013 through 2025, plus April and July 2020 supplements fielded on the 2019 panel. Each package is self-contained - questionnaire responses, the weight system and a codebook documenting every retained variable's wording, type, range and labeled values.
What changed across the SHED archive between 2013 and 2025?
Documented, dated changes: occupation and industry variables arrived in 2014, state identifiers entered the files with the 2018 wave, a dedicated Asian category joined the race variable in 2022 and was corrected in 2026, and a unique respondent identifier (shedid) first appears in the 2025 file. Each wave's codebook records its own layout.
Can SHED respondents be followed across years?
Partially. Prior-year case identifiers were added to the 2018 and 2019 files specifically to enable cross-year linking, and every recent wave ships panel weights covering the longitudinal subsample - so genuine panel inference is possible, but only across consecutive years where the identifiers line up.
How large is a SHED wave?
About 12,900 to 13,000 adult respondents in recent years - 12,934 records in the 2025 file - each carrying 815 columns in the current instrument and proportionally fewer in early waves, which predate the occupation, state and race-category additions. Accumulated across thirteen waves the archive exceeds 160,000 respondent records.
Are SHED estimates nationally representative?
Yes, once weighted. Each record carries an analysis weight, a post-stratification weight scaled to the U.S. adult population, and panel counterparts of both for the longitudinal subsample. Weighted shares reproduce the figures in the Fed's published reports; unweighted counts remain available for sample audits.
Can Datadory scope a pull to specific waves, states or question modules?
Yes. Name the wave years, the states and the modules you need - banking status, the emergency-expense series, credit access, housing, employment and gig work - and the extract returns respondent rows keyed consistently, with the matching codebook documentation attached so every labeled value arrives decoded.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.