Datadory notebook
Is HealthData.gov data free to use?
Yes - the flagship federal files clear commercial reuse outright - and Datadory delivers what those rights attach to. Datadory delivers health care services data covering the HealthData.gov estate: at least 10,000 HHS datasets spanning hospital capacity and utilization, COVID-19 patient impact, community health statistics and Medicare and Medicaid program files, typed and delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
What does the HealthData.gov estate actually hold?
Two depths share one front door, and they solve different jobs.
The catalog layer. HealthData.gov Catalog spans hospital capacity and utilization, COVID-19 patient impact, community health statistics, Medicare and Medicaid program files, substance use and workforce topics, contributed by HHS agencies alongside state and local health departments. Records carry official-source provenance flags, every tabular asset rides the same publishing chassis, and six adjacent industries draw on the same spine - distributors, supplies, technology, life and health insurance, managed care, pharmaceuticals.
The facility-week panel beneath it. COVID-19 Hospital Capacity by Facility (Historical Time Series) tracks roughly 6,000 hospitals week by week from January 1, 2020 through May 3, 2024 - about 230 collection weeks, several million facility-week rows, 127 measure columns apiece: total, staffed and pediatric bed counts, ICU occupancy, confirmed and suspected inpatient census, prior-day admissions split by age band, and critical-staffing-shortage flags. A facility key matching the CMS Certification Number ties every week of a hospital's history together.
The specialists complete the shelf: Medicare Care Compare resolves roughly 40,000 certified facilities into one typed table; Open Payments holds 15,498,687 rows across 91 columns in a single program year's general payments file; HRSA shortage designations grade every area 0 to 26 down to census tract. The catalog finds the records; these stock it.
What does a delivered row look like?
One verified state-day observation ships with the catalog record, laid out exactly as documented:
# one row = one state-day, flagship hospital capacity file
state AK
inpatient_beds 1093
inpatient_beds_used 869
inpatient_beds_used_covid 3
total_adult_patients_hospitalized_confirmed_covid 6
staffed_adult_icu_bed_occupancy 65
critical_staffing_shortage_today_yes 1Read it as a system rather than a snapshot. 869 of 1,093 staffed beds were occupied - just under eighty percent utilization - on a day when three beds held confirmed-or-suspected COVID patients and exactly one hospital in the state flagged a critical staffing shortage. The measure names do the joining: inpatient_beds against inpatient_beds_used yields utilization by simple division, and the companion *_coverage columns record how many hospitals reported each metric, so thin-reporting days get down-weighted instead of silently averaged in.
At facility grain the row family widens: identifiers - facility key, name, address, county FIPS code, subtype, metro flag - sit beneath the measures, with a documented sentinel marking suppressed values so masking never masquerades as a quiet week.
How far does coverage run, and at what grain?
Geography - the United States at every altitude: national rollups, state-day summaries, county indicators and individual facility points with FIPS codes attached. A single extract serves a one-hospital teardown and a fifty-state benchmark without re-keying anything.
Granularity - one row per hospital per collection week at the deep cut, one row per state per day on the flagship summary, one row per certified facility in the quality universe. No sampling layers anywhere, so counts recomputed from raw rows reconcile against published totals.
One interpretation rule travels with the archive: no adjustment was made for non-reporting facilities, so values reflect reports received. Filter the sentinel, weight by the coverage columns, and the series behaves; skip either step and every rate computed inherits the reporting system's gaps. Both mechanisms ship as ordinary fields in the delivery rather than footnotes in a readme.
Which fields carry the weight?
The flagship state file documents 136 columns; eleven resolve most jobs. state keys the geography. inpatient_beds and inpatient_beds_used pair into utilization. inpatient_beds_used_covid and total_adult_patients_hospitalized_confirmed_covid split the COVID load from the baseline census. staffed_adult_icu_bed_occupancy reads as a percentage straight off the row. critical_staffing_shortage_today_yes turns staffing stress into a countable event rather than an anecdote. previous_day_admission_adult_covid_confirmed and hospital_onset_covid carry the epidemiological curve. And the *_coverage family records how many hospitals reported each corresponding metric - the denominators that turn missingness into a measurable.
Definitions travel beside every field, each with its stated caveat - an unusually honest documentation standard that names its own limits. The remaining columns fold under additional-fields-on-request rather than padding every delivery, and anything pinned down during scoping arrives verified against the dictionary rather than guessed from headers.
What can you build once the estate is on site?
Four workflows pay for themselves fastest.
Utilization benchmarks. Staffed-bed counts, ICU occupancy and shortage tallies measured one way across every state give consultants and planners defensible baselines - a market's occupancy computed from the same definitions as every other market's.
Surge modeling and pandemic retrospectives. A labeled facility-week panel spanning four years, with built-in coverage counts for weighting, is scarce supervised material for epidemic models and after-action studies that need an archive which cannot drift underneath them.
Program research and exogenous features. Medicare and Medicaid program files cited straight from the publisher, and capacity, census and shortage series dropping into forecasting features alongside claims or utilization data. Newsrooms quote the federal figures with provenance intact; quant desks use them as regressors.
Who builds on this estate?
Ranked by how directly the records answer the day job:
- Investors and quant researchers stress an operator thesis against what its hospitals actually lived through, then extend screens into trajectories using editions back to 2019. See investors & quants use cases.
- Market researchers and consultants size markets from facility denominators - beds, ownership, geography - that survive scrutiny. See market researchers use cases.
- Data scientists and ML engineers get a longitudinal, labeled panel with quality machinery in the schema itself. See data scientists use cases.
- Journalists, academics and students cite facility-specific history - which hospitals ran short of ICU beds, when, and how shortages spread metro by metro - with every claim tracing to a named hospital in a named week. See journalists & academics use cases.
- Developers and data-product builders learn one publishing chassis and extend across ten thousand assets. See developers & builders use cases.
Why get HealthData.gov data through Datadory?
Because the hard part was never the first extract - it is the tenth. Schemas are set per dataset, so depth ranges from fully documented dictionaries to little more than a title. Coverage columns decide whether a rate is real or an artifact of thin reporting. A sentinel value means one thing in the facility panel and would sail into a naive parser as data. And a deliberately frozen archive looks identical to a broken feed unless someone tells you which one you have. Each is survivable once; none is fun to re-solve in every new notebook.
Datadory normalizes before delivery: field dictionaries verified against the publishers' own documentation during the August 2026 research pass, coverage counts shipped as ordinary columns, sentinels documented and handled at load, era boundaries labeled explicitly, and provenance flags carried through so official-source rows separate from contributions without a manual pass.
Where to go next
This page covers one estate inside the twenty-three-record health care services stack - sixteen primaries plus seven pooling in from neighboring industries. Start with the health care services data hub for the full pooled view, or the ranked shortlist of the best health care services datasets.
Go deeper three ways: the HealthData.gov Catalog product page documents the field dictionary above and takes sample requests scoped to your states and topics; the facility-week capacity panel carries the four-year archive measure by measure; and the breadth-versus-precision argument continues on the HealthData.gov Catalog vs HRSA Data Warehouse comparison. The publishing house behind it maps onto the HHS source profile, and the operational counterpart across the Atlantic is NHS England Statistics & Data Collections.
When you want real rows instead of descriptions, request a sample - the field dictionary travels with it.
| Dimension | Coverage |
|---|---|
| Geographic | United States at every altitude - national rollups, state-day summaries, county indicators and individual facility points carrying FIPS county codes; shortage designations geocode from facility points through census tracts to state rollups |
| Granularity | One row per hospital per collection week (127 measure columns) at the deep cut; one row per state per day on the flagship summary (136 documented columns); one row per certified facility in the roughly 40,000-row quality universe |
| Record | What it measures | Grain | Where it fits |
|---|---|---|---|
| HealthData.gov Catalog | Discovery over American health publishing: capacity and utilization, community health, Medicare and Medicaid program files, substance use, workforce | At least 10,000 cataloged datasets, roughly 1,269 answering 'hospital'; quality 9/10 | Index and deep cut at once - one request serves a national dashboard or a single-market teardown |
| COVID-19 Hospital Capacity by Facility (Historical Time Series) | Facility-level utilization through the pandemic: beds, ICU occupancy, inpatient census, admissions, staffing shortages | About 6,000 hospitals x 230 collection weeks, 127 measure columns per row | The four-year labeled archive that surge models and after-action studies train on |
| Medicare Care Compare (Hospitals, Nursing Homes, Home Health, Dialysis) | Certified-provider quality: star ratings, inspection composites, staffing, penalties, ownership | Roughly 40,000 facilities keyed by CMS Certification Number; editions back to 2019 | The join spine - any American provider file attaches to it on CCN |
| Open Payments (Physician Payments Sunshine Act) | Industry payments and ownership interests reported for US clinicians and teaching hospitals | 15,498,687 rows x 91 columns in the 2024 general payments file; program years 2019-2025 | Row-grain evidence on the influence economy of American medicine |
| HRSA Data Warehouse - Health Workforce Shortage Areas | Official HPSA and MUA/P designations with National Health Service Corps scores | 21 documented fields, geocoded from facility points through census tracts to states; quality 9/10 | Access geography for site selection, eligibility screening and research |
| WHO Global Health Observatory Data Repository | Global indicators: workforce density, hospital beds, service coverage, mortality, financing | 3,093 indicator codes across 194 Member States with sex and age splits preserved; quality 9/10 | The cross-country panel for when the commercial question leaves US borders |
| NHS England Statistics & Data Collections | Operational statistics of the English NHS: referral-to-treatment waits, A&E attendances, bed occupancy | About 40 work areas; RTT monthly since March 2007; KH03 occupancy back to 1987-88 | Trust-by-specialty waiting-list pressure no international panel resolves |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
HealthData.gov Catalog
COVID-19 Hospital Capacity by Facility (Historical Time Series)
Medicare Care Compare: Facility-Level Quality Data for Every Certified Provider
facility_id · hospital_type · hospital_ownership …+5 more
Open Payments (Physician Payments Sunshine Act)
HRSA Data Warehouse - Health Workforce Shortage Areas
KFF State Health Facts & Health Costs Data
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What comes in a Datadory HealthData.gov delivery?
Typed rows in one documented schema: the state-day capacity summary with its eleven core attributes, or the facility-week panel's identifiers and measure columns, with coverage counts, sentinel handling and provenance flags shipped as ordinary fields. The sample arrives shaped identically to the production feed, field dictionary attached.
Why does the hospital capacity series end in 2024?
Reporting stopped being mandatory after May 3, 2024, when hospitals were no longer required to submit COVID-19 admission figures, so the daily state run ends there while voluntary surveillance continued. Nothing dated later will appear, and that fixity is the feature - backtests and retrospectives get an archive that cannot drift underneath them.
How many datasets does the HealthData.gov estate cover?
At least 10,000 cataloged assets - the discovery layer caps its own result set at 10,000, so the true total may run higher - and roughly 1,269 answer the query 'hospital' alone. Size the shelf by topic rather than by the grand total; a sample can cut the catalog to a single theme.
Can a delivery be scoped to specific states, facilities or subjects?
Yes. Name the states, counties, facilities, programs or subject areas when you request the sample and it ships filtered to that scope in the documented field shape. The ongoing feed follows the same structure, delivered by API, scheduled files, or straight into your warehouse, daily, weekly, or hourly.