Datadory notebook
How Data Scientists Use Consumer Finance Data
Data scientists use consumer finance data to train credit-risk models, engineer household-wealth features, monitor the credit cycle, and benchmark servicer conduct - from free regulator sources: the CFPB's ~17.24 million complaints, Federal Reserve G.19 series back to January 1943, SCF microdata, HMDA loan-level mortgages and Global Findex across 141 economies. All 24 datasets Datadory catalogs for the industry are free.
1,744 datasets. Pick your catch.
How do data scientists actually use consumer finance data?
Four jobs dominate the desk. Train default models on loan-level repayment outcomes (LendingClub's mirror). Engineer household-balance-sheet features (SCF, SHED, FDIC). Track the aggregate credit cycle monthly (G.19). Mine complaint text for conduct signals (CFPB). Everything else - fair-lending screens on HMDA, financial-inclusion sizing with Findex, retail nowcasting with Census EITS - hangs off those four.
The ranked shortlist below is ordered by Datadory quality score as of August 2026.
Which dataset anchors a default-model training set?
Start with the LendingClub Loan Data (Kaggle Mirror): about 2.26 million accepted personal loans from June 2007 through 2018 Q4 plus roughly 2.76 million rejected applications, across 151 documented fields that include monthly performance and settlement outcomes. It is declared commercial delivery terms, in a ~1.36 GB zipped package that needs Kaggle sign-in but no payment.
For macro covariates, join the Federal Reserve G.19 Consumer Credit release: 118 series totaling about 78,000 monthly observations, reaching January 1943, with auto, card and personal-loan rate series beginning as early as June 1971. Origination-vintage features line up cleanly against G.19 months, and the whole machine-readable package is under 6 MB zipped in SDMX form.
Where does household wealth and resilience microdata come from?
Wealth modeling runs on the Federal Reserve Survey of Consumer Finances (SCF): triennial waves 1989 through 2022 covering family assets, liabilities, income and pensions. The 2022 public extract holds 22,975 records - 4,595 families times five multiply-imputed implicates - published under DOI 10.17016/8799 with no registration wall. The non-negotiable step is applying the published replicate weights; the Fed's own Standard Error Documentation exists because ignoring multiple imputation and the complex design yields incorrect standard errors.
Segmentation work adds the FDIC National Survey of Unbanked and Underbanked Households: roughly 30,000 households per biennial wave across eight waves 2009-2023, with 567 documented variables in the 2023 metadata and an hhmultiyears.zip spanning all waves plus weights. Cross-country extension uses the World Bank's Global Findex, covered below.
How do you mine complaint text without hitting rate limits?
The CFPB Consumer Complaint Database indexes roughly 17.24 million complaints - Socrata's record count stood at 17,242,643 as of 2026-08-21 - one row each from December 1, 2011 onward, refreshed generally daily under commercial delivery terms, with product and issue taxonomies, company responses and opt-in narratives in a ~1.42 GB bulk ZIP.
For pipelines rather than one-off pulls, the companion CFPB Consumer Complaint Database API (ccdb) needs no integration key: field-level filters, aggregations, JSON or CSV output, page sizes of 1-100 records, and search_after for deep pagination past what offset paging allows. A daily incremental job filters on the date field, pulls new rows, and appends - no authentication plumbing at all.
Modeling caveats are printed on the source itself: the Bureau states the database is not a statistical sample of consumers' experiences, low complaint volume does not necessarily mean little harm, and narratives are not verified. Treat complaint counts as conduct telemetry skewed toward products with dissatisfied, articulate customers - a useful signal, not a population estimate.
What does a fair-lending screen need from HMDA?
Loan-level mortgage work runs on two complementary sources. The CFPB Home Mortgage Disclosure Act (HMDA) Data publishes bulk files from 2007 through filing year 2025 - one row per application with lender LEI, rate spread, applicant demographics and action taken, down to census tract, privacy-modified before release. Nationwide first-lien owner-occupied volume ran ~4.8M-8.3M rows per year across 2007-2017, and the 2023 dynamic LAR ZIP alone is ~631 MB compressed spanning millions of rows across roughly 100 fields.
Can you extend the stack internationally?
The World Bank Microdata Library - Global Findex Catalog lists 717 Findex study records with questionnaires and dictionaries across the five waves - useful when you need instrument wording to align indicators across 2011-2024.
Which datasets earn a slot on the ranked shortlist?
Ranked by Datadory quality score, then scale, these are the records a working consumer-finance stack actually loads:
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
CFPB Consumer Complaint Database
CFPB Consumer Complaint Database API (ccdb)
Federal Reserve G.19 Consumer Credit Data
Eight documented core fields · with holder-level
LendingClub Loan Data (Kaggle Mirror)
Federal Reserve Survey of Consumer Finances (SCF)
Federal Reserve Survey of Household Economics and Decisionmaking (SHED)
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Where can I get household wealth microdata with correct standard errors?
The Federal Reserve Survey of Consumer Finances: triennial waves 1989-2022 with a 2022 extract of 22,975 multiply-imputed records for 4,595 families, downloadable free under DOI 10.17016/8799. Apply the published replicate weights and follow the Standard Error Documentation - ignoring imputation and sample design yields incorrect standard errors.
Which time series tracks the US household credit cycle?
The Federal Reserve G.19 Consumer Credit release: 118 series totaling roughly 78,000 monthly observations on revolving versus nonrevolving credit by holder type, running from January 1943 to June 2026, with auto, card and personal-loan rate series beginning as early as June 1971 in CSV and SDMX packages.