Datadory notebook

HCUP NIS discharge data: what a purchase should put in your warehouse

Datadory delivers the health care REITs data an HCUP NIS discharge data purchase is really after: about 7 million all-payer inpatient stays a year weighted past 35 million hospitalizations, every row carrying expected payer, diagnoses, procedures, charges and disposition back to 1988, beside complete state inpatient universes across 48 partner states plus DC - typed rows delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What does an HCUP NIS discharge data purchase actually buy?

The query says purchase; the substance is the only all-payer encounter ledger American hospital care has. [{n}]({p}) is AHRQ's eight-database family built from the near universe of encounter-level hospital care, and its flagship National (Nationwide) Inpatient Sample carries about 7 million discharges per year, each holding more than 100 data elements, discharge-weighted to represent more than 35 million hospitalizations nationally. Inpatient records run from 1988 onward, which makes this the deepest consistent hospitalization series available anywhere for longitudinal work.

Stated plainly, here is what the money is for: PAY1 codes who was expected to pay for each stay - Medicare, Medicaid, private insurance, self-pay, no charge - and no single-program federal file observes that column, because each one sees only its own claims. Underwriting a skilled-nursing operator whose census skews Medicaid, sizing senior-housing demand in a referral market, benchmarking post-acute flow downstream of a specific hospital: all of it starts from that one field sitting on the same row as the diagnoses and the outcome.

Position inside the slice matters as much as scale. The health care REITs pool stacks its records in layers - facility grading, then charge-per-service-line, then this one, the only layer where each stay keeps its payer, its conditions and how the episode ended together. That singularity earns the 8/10 quality score against a catalog mean of 7.81 across 1,744 datasets.

Most of the shopping effort this query attracts online goes to paperwork and file formats rather than analysis. Get a sample of NIS discharges cut to named databases, states, years and fields instead, and judge the records before committing to a feed.

What does one discharge row look like?

One row per de-identified hospital stay, flat in the shape it lands:

# one row = one de-identified inpatient stay -- illustrative, NIS-shaped

KEY_NIS      : <encounter key>          HOSP_NIS    : <masked hospital id>
AGE          : <years>                  FEMALE      : <sex indicator>
PAY1         : <1 Medicare | 2 Medicaid | 3 private | 4 self-pay | 5 no charge>
DXCCSR1..N   : <all-listed diagnoses -> CCSR categories, ICD-10-CM>
PRCCSR1..N   : <all-listed procedures -> CCSR categories, ICD-10-PCS>
TOTCHG       : <total charges for the stay, nominal dollars>
DISPUNIFORM  : <routine | transfer | home health | died>
DISCWT       : <discharge weight expanding the sample to national totals>

Read the anatomy rather than the placeholders. Six columns settle most utilization questions: PAY1 names who was expected to pay - the split program-specific claims never see - DXCCSR and PRCCSR collapse raw ICD-10 codes into analyzable condition and procedure categories, TOTCHG prices the stay, and DISPUNIFORM records how the episode ended, which is where post-acute demand lives. De-identification generalizes dates to quarter and masks hospital identity while keeping it linkable: enough resolution to model a market, not enough to re-identify a patient. A requested sample replaces every bracket with real records.

Which fields carry the payer-mix analysis?

PAY1 codes the expected primary payer across Medicare, Medicaid, private insurance, self-pay and no charge. It is the reason this record exists as a product at all: revenue quality at a tenant operator is a payer question before it is a census question, and here the mix sits measurable per hospital service area rather than asserted in a management deck.

The rest of a discharge runs past 100 elements - admission origin, length-of-stay flags, the Chronic Condition Indicator and Elixhauser comorbidity flags among them - and travels folded into the request step rather than cluttering the core dictionary. Field names follow documented coding conventions; confirmed column names are pinned against live records when your sample ships.

How far does coverage reach, and at what grain?

Three boundaries shape what a discharge panel can honestly claim, and all three are knowable before any commitment.

Geography splits two ways. Five nationwide samples produce national estimates; three state families - SID, SEDD and SASD - carry the complete discharge universe of every nonfederal acute care hospital in 48 partner states plus DC. A market-level question counts every discharge in the states you name; a national trend question weights a sample up to the country.

Time runs deep with one seam. Inpatient coverage starts in 1988, and the National Inpatient Sample was redesigned in 2012, which changes how sampling weights behave. Pool pre- and post-2012 years without handling the break and your trend line quietly changes meaning halfway through. The Kids' Inpatient Database ships as its own edition every three years, so pediatric trend work plans around edition dates rather than calendar years.

Grain is the stay, not the person. Dates arrive generalized to quarter, so short-stay timing analyses lose precision while annual and quarterly aggregations lose nothing. Read that as a design property rather than a defect: it is what lets a market model run without touching patient identity.

Which database answers your question - national sample or state universe?

The family divides by geography and service line, and picking the right half is most of the analysis design. Eight databases, two halves:

Two examples make the choice concrete. Counting joint replacements in the three states where a portfolio operates is SID work - the state universe counts every discharge, no weighting error attached. Measuring how national Medicaid stays moved over a decade is NIS work - the weighted sample is the defensible estimator. Ask for both when the question spans them; the samples return shaped to the databases named.

What can you build once the discharges are on site?

Four builds recur across the buyers who request this record:

Which datasets pair with NIS discharges?

The encounter ledger gains meaning once identity, price and place sit beside it, and the same slice supplies all three:

  • CMS Medicare Provider Charge Data (Inpatient & Outpatient) is the pricing counterpart: one row per hospital by DRG for inpatient service years 2013 through 2024 and by APC for outpatient 2015 through 2024, across 3,000+ IPPS hospitals keyed by CCN. It reflects Original Medicare fee-for-service claims only - Medicare Advantage stays are absent entirely - which is precisely the gap PAY1 fills. Pairing them is the standard move: CMS rows give hospital-level price and volume benchmarks, NIS encounters give the all-payer denominator.
  • Yahoo Finance Health Care REIT Quote Pages closes the loop from health systems to listed landlords - daily OHLCV history for WELL, VTR, ARE, DOC, OHI and CTRE - so utilization shifts test against the equity tape rather than a thesis.

Who buys all-payer discharge data?

Ranked by how directly the record answers the day job:

Why buy NIS discharges delivered rather than assembled?

Start with a sample: name the databases, partner states, years and fields, and the extract returns cut to exactly that slice with the field dictionary attached - shaped identically to the standing feed, so anything prototyped survives into production unchanged.

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Health Care REITs United States: national estimates from five nationwide…

HCUP - Healthcare Cost and Utilization Project

PAY1 · HOSP_NIS · DISCWT

Health Care REITs United States

CMS Medicare Provider Charge Data (Inpatient & Outpatient)

Health Care REITs United States - national, state, county/HRR and facility level…

CMS Data Hub - Medicare & Medicaid Datasets

Health Care REITs United States at four levels: hospital referral regions (about…

Dartmouth Atlas of Health Care

Cohort · Eventname · Event_label …+1 more

Health Care REITs United States (NYSE and Nasdaq listings)

Yahoo Finance Health Care REIT Quote Pages

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

What arrives in an HCUP NIS discharge delivery?

Typed rows rather than fixed-width archives: one de-identified stay per row carrying AGE, FEMALE, PAY1, CCSR-grouped diagnoses and procedures, TOTCHG, DISPUNIFORM, a masked hospital identifier and the discharge weight, with the SID, SEDD and SASD state universes labeled alongside the nationwide samples. Name the databases, states and years when requesting a sample and it returns shaped that way.

Which HCUP database should a market-level question use?

State universes for markets, national samples for trends. SID counts every inpatient discharge in each of 48 partner states plus DC, so a question about named states carries no weighting assumptions; NIS, NEDS, NASS, KID and NRD are designed samples that weight up to national estimates. Both families can sit in one delivery under the same encounter spine.

How does the 2012 NIS redesign affect multi-year work?

The redesign changed how the sample is drawn and therefore how discharge weights behave, so naively pooled pre- and post-2012 series inherit a structural break at the boundary. Handle it with explicit year-segment treatment and the weights still support defensible national estimates across the whole span - weights ship named and typed per vintage so the arithmetic stays auditable.

Can a delivery be scoped to my states, years and databases?

Yes. Name the databases, the partner states, the year range and the fields you need - one state's inpatient universe for a decade, national-sample-only estimates, emergency visits for selected metros - and the sample returns shaped to exactly that slice with the field dictionary attached. The standing feed keeps the identical shape.