Datadory notebook

How Data Scientists Use Health Care Facilities Data

Data scientists use health care facilities data as three joined layers: HealthData.gov's 127-column facility-week hospital capacity panel for utilization dynamics, HRSA's AHRF and Uniform Data System files for county supply and per-awardee outcomes back to 1999, and CMS's NPPES NPI Registry as the entity-resolution spine. Of the 13 primary datasets Datadory catalogs for this industry, 9 are free.

1,744 datasets. Pick your catch.

What does a health care facilities pipeline actually look like?

Three facts frame every design decision below. First, the deepest time series ended on purpose - HHS stopped collecting hospital capacity data after May 3, 2024. Second, the widest table is county-level: the AHRF facilities file alone carries 716 columns. Third, only one source here enumerates providers individually - the NPI Registry described below - and it is the join key for everything else.

Which panel should anchor your utilization model?

HealthData.gov - COVID-19 Hospital Capacity by Facility (Historical Time Series) is the only facility-week utilization panel in the slice, and structurally it behaves like a benchmark ML dataset: roughly 6,000 hospitals across about 230 collection weeks from January 1, 2020 through May 3, 2024, several million rows total, each carrying 127 columns. Identifiers pin every row to a facility - hospital_pk matching the CCN where one exists, plus name, address, fips_code, hospital_subtype and an is_metro_micro flag - while measures arrive with three suffixes that tell you how to aggregate them: _avg for 7-day averages, _sum for weekly totals, _coverage for how many times the facility reported that element during the week.

Two caveats belong in your feature-store docs. No statistical adjustment was made for non-reporting or missing facilities - values reflect only reports received - so the _coverage columns double as a per-field missingness indicator you get for free. And the series is static: reporting ended May 3, 2024, ideal for retrospective surge studies, useless as a live signal.

Where do county-level provider-supply features come from?

Two sibling file families extend it below the county line. Health Center Service Delivery and Look-Alike Sites ships one row per geocoded site across 56 columns, including FQHC Site NPI Number - your bridge to the provider master described later. HPSA and MUA/P designation files add shortage geometry: primary care HPSA detail runs about 48 MB, MUA/P detail about 11 MB, with polygon boundaries and centroid points in SHP form. The interactive layer of the same warehouse is cataloged separately as HRSA Data Warehouse - Health Center & Facility Finders.

Everything lands over plain HTTP with no registration under public-domain terms whose cards read "Usage limitations: None". One caveat: UDS-derived files carry a FOIA notice stating any alteration or format conversion is the user's sole responsibility.

Can you build a 27-year outcome panel from UDS filings?

Yes, and it is the quiet advantage of this industry: standardized annual reporting for the same organizational units going back to 1999. The HRSA Health Center Program Uniform Data System (UDS) Data 2025 release centers on the H80 awardee workbook - about 28.3 MB of XLSX across 38 sheets covering 1,356 health centers, with Tables 3A through 9E holding patients, visits, staffing, quality measures and financials. A look-alike workbook (~2.9 MB) and a service-delivery-sites CSV (~12.8 MB, 56 columns) round out the set, all commercial delivery terms.

Feature-engineering notes that save a day: measure columns arrive coded, T3a_L1_Ca-style, keyed to the UDS Tables manual rather than self-documenting, so build a code dictionary before writing transforms. Blank cells can mean suppression rather than zero - edit checks are logged in the Edits sheet. And the featured year lags: the 2025 Coversheet stamps DateOfLastReportRefresh 05/07/2026.

How do you resolve facilities into one entity spine?

Respect the method condition CMS attaches: no charge, but per-hour limits apply and bulk queries must use the DDS file instead of the API. The API caps requests at 200 results with skip limited to 1,000 - roughly 1,200 records over six calls - marking where interactive lookups end and batch ingestion begins.

For geometry, the companion portal is HRSA Data Warehouse - Health Center & Facility Finders: an unauthenticated JSON locator takes lat/lon/radius queries, registered web services need a free token tied to a declared calling domain, and facility locator records refresh daily - same-day record dates were observed at research time. A practical spine: land the monthly NPPES file as master, sync HRSA points daily, then left-join AHRF county features and UDS measures on NPI and site name, diffing basic.last_updated for change detection.

How do you assemble the stack step by step?

A numbered walkthrough from empty cluster to joined analytical table:

Every step runs on free, public-domain inputs - the stack fits a research budget of zero.

Which datasets make the ranked shortlist?

Ranked by Datadory's quality score as of August 2026, weighted toward what a modeling workflow can ingest without procurement:

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Health Care Facilities United States

COVID-19 Hospital Capacity by Facility (Historical Time Series)

Health Care Facilities United States including territories

HRSA Data Downloads - AHRF, Health Center Sites, HPSA & Transplant Files

Health Care Facilities United States and territories - 50 states

HRSA Data Warehouse - Health Center & Facility Finders

Distance

Health Care Facilities United States - all 50 states

HRSA Health Center Program Uniform Data System (UDS) Data 2025

Health Care Facilities United States and its territories - one record for every…

NPPES NPI Registry - National Provider & Facility Identifier Search

number

Health Care Facilities United States - all 50 states

American Hospital Association Organizational Hub & Data Insights

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Is health care facilities data free for machine learning projects?

Mostly. Nine of the 13 primary datasets Datadory catalogs for health care facilities are free, including the HealthData.gov capacity panel, both HRSA records and the NPPES NPI registry - all commercial delivery terms or freely releasable. Only the two American Hospital Directory services and Definitive Healthcare charge.

What is the best dataset for modeling hospital capacity?

HealthData.gov's COVID-19 Reported Patient Impact and Hospital Capacity by Facility series: one row per hospital per collection week across roughly 6,000 hospitals and about 230 weeks, with 127 columns spanning ICU occupancy, admissions and staffing shortages from January 1, 2020 through May 3, 2024 via the Socrata API.

How do I deduplicate provider and facility records?

Key everything on the 10-digit NPI from CMS's NPPES registry. The Read API v2.1 refreshes daily for lookups capped at 200 results, while bulk consumers must take the monthly dissemination ZIP of about 1,098 MB compressed plus weekly incremental and deactivation files.

Where can I get FQHC performance data for benchmarking?

HRSA's Uniform Data System: the 2025 H80 workbook covers 1,356 health center awardees across 38 sheets with patients, visits, staffing and quality measures, and archived editions reach back to 1999 through the Electronic Reading Room - all commercial delivery terms in XLSX and CSV.