Datadory notebook

Household Wealth Microdata With Replicate Weights Data: Dataset Structure and Field Coverage

Datadory delivers household wealth microdata with replicate weights data covering comprehensive field definitions, entity mappings, and historical time series — structured for direct analytics and delivered on demand.

1,744 datasets. Pick your catch.

What counts as household wealth microdata?

Household wealth microdata is one row per family or respondent carrying balance-sheet detail: assets on one side, debts on the other, plus the income, pensions and demographics needed to cut distributions. In Datadory's consumer-finance slice, three records deliver that grain: the Federal Reserve Survey of Consumer Finances (SCF), the Federal Reserve Survey of Household Economics and Decisionmaking (SHED), and the World Bank Global Findex Database 2025. The rest of the 24-dataset slice sits elsewhere - the G.19 publishes aggregates, not families, and HMDA is loan-level rather than household-level.

The "replicate weights" half of the query separates a defensible estimate from a wrong one. The Board's guidance for the SCF warns that ignoring multiple imputation and the complex sample design yields incorrect standard errors. Weights (the SCF's WGT, SHED's post-stratification weights, FDIC's hhsupwgt) fix point estimates; replicate weights are what make the variance around them honest.

Why does the SCF ship five records per family?

Because missing answers are multiply imputed five times. Every family appears as five records - implicates - so the 2022 public extract totals 22,975 rows for 4,595 interviewed families after disclosure-avoidance removals. The YY1 column identifies which implicate a row belongs to and Y1 carries the family identifier within the survey year, which is how you reassemble them.

How do you apply the replicate weights correctly?

The weights arrive in files separate from the data, and treating them as optional is the most common mistake with this dataset. Four steps cover a standard analysis:

Skip steps 3 and 4 and your confidence intervals come out too narrow; the Board states plainly that ignoring multiple imputation and the complex sample design yields incorrect standard errors.

How does SHED complement the SCF between waves?

Grain details matter when joining the two. SHED covers individual adults - 12,934 respondents by 815 variables in the 2025 public CSV, 47.8 MB unzipped - representable via post-stratification weights, with state identifiers (ppstaten) included from the 2018 wave onward and nothing finer than state published. Formats narrowed in February 2026 when SAS distribution was discontinued, leaving CSV and Stata with the Fed confirming identical content. Like the SCF it is commercial delivery terms, published under DOI doi.org/10.17016/8960.2 for the 2025 report.

Can any source take household wealth analysis global?

Only one comes close, and it changes what "wealth" means. The World Bank Global Findex Database 2025 surveyed about 144,090 adults across 141 economies across five waves since 2011 (2011, 2014, 2017, 2021, 2024). The country-level file carries 438 columns under commercial delivery terms 4.0; respondent-level microdata runs 199 columns and about 144,090 rows in a 50.4 MB unzipped CSV, gated behind Microdata Library research-use terms (survey reference WLD_2024_FINDEX_v02_M, DOI ): statistical or scientific research only, no redistribution without written agreement.

Findex measures account ownership, mobile money, savings behavior and credit access - demand-side inclusion - not net worth, so it frames cross-country context around balance-sheet work rather than substituting for it. Disaggregation by gender, income quintile, age, education, labor force status and rural/urban status is built into the microdata. For US-only wealth questions the SCF remains unmatched; between triennial SCF waves, SHED keeps annual resolution.

Which survey fits which wealth question?

A practical pattern falls out of the table: anchor distribution claims on the SCF, refresh resilience and sentiment measures annually with SHED, reach for G.19 series when the question turns to aggregate balances rather than families, and use FDIC microdata when banking status needs state or MSA cuts.

Who works with weighted wealth microdata day to day?

Four workflows dominate. Academic and policy researchers computing inequality statistics - concentration measures, median-versus-mean gaps by age band - are canonical users, and they are exactly who the replicate-weight discipline protects: a misstated standard error survives peer review badly. Journalists quoting wealth gaps need the same treatment, since a headline number without its variance band invites challenge.

Verified SCF 2022 download package
FileContentsSize
scfp2022excel.zipSummary Extract: 357 columns x 22,975 rows (SCFP2022.csv)~2.2 MB
scf2022ascii.zipFull Public Data Set, ASCII with variable maps~38 MB zipped
scf2022s.zipFull Public Data Set, Stata~9 MB zipped
scf2022rw1*.zipReplicate weight files keyed on X42001~26 MB zipped
codebk2022.txtCodebook documenting every field including the weights~490 KB

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Consumer Finance United States - nationally representative of families

Federal Reserve Survey of Consumer Finances (SCF)

Consumer Finance United States, nationally representative of adults

SHED Public Use Data Files (2013-2025)

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Do I really need the SCF replicate weights?

Yes, whenever you report a variance measure. The Board's Standard Error Documentation warns that ignoring multiple imputation and the complex sample design yields incorrect standard errors, and each family contributes five implicates that must be combined. Point estimates need the WGT weight alone; confidence intervals need the replicates.

How many records are in the SCF 2022 extract?

22,975 records covering 4,595 interviewed families, because missing answers are multiply imputed five times. YY1 identifies which implicate a row belongs to and Y1 identifies the family within the survey year. The Summary Extract carries 357 variables while the Full Public Data Set reaches roughly 5,300 per family.