MIMIC-IV - Medical Information Mart for Intensive Care
Datadory delivers Health Care Technology data covering MIMIC-IV - Medical Information Mart for Intensive Care: 364,627 de-identified patients from a single US academic medical center between 2008 and 2022, across 546,028 hospitalizations and 94,458 ICU stays with hour-level vitals, labs, medications, diagnoses, discharge and radiology notes, plus 377,110 chest X-rays. Delivered daily, weekly, or hourly.
What is the MIMIC-IV - Medical Information Mart for Intensive Care dataset?
One academic medical center's two-decade intensive care record, opened up as structured data instead of filing cabinets. MIMIC-IV - Medical Information Mart for Intensive Care assembles the de-identified electronic health record of patients treated at Beth Israel Deaconess Medical Center in Boston between 2008 and 2022, curated by the MIT Laboratory for Computational Physiology and distributed under the PhysioNet banner. Version 3.1 (published October 2024) covers 364,627 unique individuals. The hosp module spans hospital-wide ground for 546,028 hospitalizations serving 223,452 of those people: admissions and transfers, laboratory results, microbiology cultures, prescriptions, electronic medication administrations, pharmacy orders, ICD-coded diagnoses and procedures, billing and service assignments. The icu module follows 65,366 individuals through 94,458 ICU stays with hour-level charted vitals, fluid inputs and outputs drawn from the bedside clinical system. An ed module adds emergency department visits with triage scores, and companion note, cxr and ecg modules carry discharge and radiology free text, 377,110 chest X-ray images across 227,835 studies, and electrocardiogram machine measurements. It scores 10/10 in our catalog - a mark only 145 of the 1,744 datasets we researched reach.
What do sample rows look like?
One patient-to-stay chain, exactly as it lands in your warehouse. Values shown are the documented example values carried on each field definition:
subject_id : 10000032
hadm_id : 22595853 stay_id : 31933343
admission_type : EW EMER.
admittime / dischtime : shifted timestamps inside a 2100-2200 window
anchor_year / anchor_age : populated hospital_expire_flag : 0 or 1
itemid : populated valuenum / valueuom : result plus unit
charttime / storetime : event time versus documentation time
starttime / endtime : drug administration window
icd_code / icd_version : populated (9 or 10)
insurance / language /
marital_status / race : populated per hospitalizationRead it as a join kit. subject_id keys the person, hadm_id keys the hospitalization, stay_id keys the ICU or ED stay, and itemid joins every event row back to the lab and charted-variable dictionaries. The anchor-year pair restores calendar arithmetic on top of the shifted timestamps, so a timeline computed in 2153 still means exactly what it says. Sepsis windows, dosing intervals and readmission gaps all hang off that identifier spine.
What fields does the dataset include?
Thirteen documented field groups anchor the dictionary, verified against the source's published schema documentation during research. Module-specific detail tables - microbiology antibiograms, pharmacy fill lines, ED triage and med-reconciliation rows, radiology report text and imaging metadata - ride on the same identifiers and ship under additional fields scoped at sample request rather than promised in every cut.
What does coverage look like across geography, time and granularity?
Geography - deliberately deep rather than broad: a single US academic medical center, Beth Israel Deaconess Medical Center in Boston, Massachusetts. Every lab draw, charted vital, medication dose, note and image comes from one institution's real operations, which is precisely why the record supports internal-validity claims most multi-site aggregates cannot.
Temporal - 2008 through 2022 for the core record, with the chest X-ray subset spanning 2011 to 2016. Raw calendar dates are replaced during de-identification: patient events land inside a 2100-2200 window, and each patient carries an anchor year and anchor age so year-level aging, seasonality and multi-admission spacing stay computable without exposing anyone's true dates.
Granularity - four nested levels: patient, then hospitalization, then ICU or ED stay, then the individual event row underneath - one row per lab result, per charted observation, per medication administration. Imaging sits at DICOM study level, paired with its radiology report.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Hourly suits live evaluation rigs - a model under test reading the same stay-level panels your researchers do, refreshed fast enough that a demo dashboard never shows stale vitals. Daily suits production pipelines: cohort features recomputed each morning while the previous day's runs reconcile cleanly against the last. Weekly suits research rhythms - benchmark reruns, paper revisions, quarterly cohort refreshes where stability matters more than latency.
Every cut reuses the identical field names and the same identifier spine, so snapshots concatenate into longitudinal panels without remapping a single column. Request the modules you need - hospital-wide, ICU, emergency, notes, imaging - and the sample arrives cut to them, with the field dictionary and coverage profile attached.
Who uses this data, and for what?
- Sepsis, deterioration and outcome modeling - hour-level charted vitals joined to lab results and medication administrations give onset-prediction models a defensible ground truth; mapped to the ml-model-training use case.
- Clinical language models - discharge summaries and radiology reports aligned to structured outcomes turn free-text benchmarks into measurable ones, from summarization fidelity to diagnosis-code prediction.
- Imaging-and-report pairs - 377,110 chest X-ray studies keyed to their radiology text support multimodal retrieval, report-generation and label-noise studies at a scale few institutional partners will share.
- Operations and health-services research - admission types, transfer chains, service assignments and payer mix support length-of-stay, readmission and capacity studies grounded in actual hospital operations.
- Pharmacovigilance and dosing studies - start-and-stop timestamps on drug administrations plus fluid inputs and outputs expose treatment timing at minute resolution.
- Benchmark lineage - the record sits behind a large share of published critical-care ML work, so a new model's numbers land on a comparable footing; see the citation-grade research use case.
Which personas get the most value?
Data scientists and ML engineers get the deepest single-site clinical panel in the catalog: event-level grain, stable identifiers and a dictionary that documents itself, the difference between a publishable cohort and a weekend of schema archaeology. Journalists and academics get citable, versioned numbers on what intensive care actually looks like - stays, mortality flags, payer mix - without waiting on an institutional partnership. Investors and quants read care patterns, utilization and documentation intensity as ground truth beneath health-technology theses pitched on operational efficiency. Product teams building healthcare software get a realistic test corpus whose quirks - missingness, shift charts, unit mismatches - match production reality far better than synthetic fixtures ever will.
What should I know before requesting a sample?
Three things worth knowing upfront. First, scope is depth over breadth: one Boston academic medical center, so treat external validity as a hypothesis your own multi-site data tests rather than something this record promises. Second, time is de-identified - absolute calendar dates are shifted into a 2100-2200 window, so analyses needing true wall-clock dates map through the anchor-year pair instead; plan for that join before you build. Third, volume needs scoping: the imaging and waveform companions are measured in hundreds of thousands of studies, so name the modalities and year windows you want when requesting a sample and the cut arrives sized to the question rather than to the archive.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
subject_id | integer | Patient identifier; one row per patient in the patients table, repeated across admissions and event tables. | 10000032 |
hadm_id | integer | Unique random integer per hospitalization (range 2000000-2999999), defined in the hosp admissions table. | 22595853 |
stay_id | integer | Identifier for an ICU or ED stay, defined in the icu icustays and ed edstays tables. | 31933343 |
admittime / dischtime | datetime | Timestamps of hospital admission and discharge in the hosp admissions table; dates shifted into a 2100-2200 window during de-identification. | |
admission_type | enum | Urgency class of the visit; nine values including ELECTIVE, URGENT, EW EMER., DIRECT EMER., OBSERVATION ADMIT, EU OBSERVATION, DIRECT OBSERVATION, AMBULATORY OBSERVATION and SURGICAL SAME DAY ADMISSION. | EW EMER. |
insurance / language / marital_status / race | string | Per-hospitalization demographic and payer fields recorded in the hosp admissions table. | |
hospital_expire_flag | boolean | Binary flag in the hosp admissions table: 1 if the patient died during the hospitalization, 0 otherwise. | |
itemid | integer | Key joining event rows to the dictionary tables for laboratory items and charted ICU variables. | |
charttime / storetime | datetime | Event time and documentation time for measurements in lab, charted-output, input and procedure event tables. | |
valuenum / valueuom | number | Numeric result and unit of measure for laboratory and charted observations. | |
starttime / endtime | datetime | Start and stop times for drug administrations and fluid or procedure inputs. | |
icd_code / icd_version | string | Diagnosis or procedure code plus ICD version (9 or 10) in the coded diagnosis and procedure tables. | |
anchor_year / anchor_age | integer | Patients-table fields giving the shifted anchor year and the patient's age at that year, enabling year-level temporal reasoning without exposing true dates. |
Questions buyers ask
How many patients does the dataset cover?
364,627 unique individuals in version 3.1. The hospital-wide module covers 546,028 hospitalizations serving 223,452 of them, the ICU module tracks 65,366 individuals through 94,458 stays, and the imaging companion adds 377,110 chest X-ray images across 227,835 studies. Counts that size make subcohort work - rare diagnoses, specific age bands - statistically routine.
Which fields identify a patient across modules?
Three nested identifiers: subject_id keys the person, hadm_id the hospitalization, and stay_id the ICU or ED stay. itemid then joins event rows to the dictionary tables that define each laboratory item and charted variable. Because every module reuses that spine, a hospital-wide admission, its ICU hours and its notes attach without fuzzy matching.
How are dates handled in a de-identified dataset?
HIPAA Safe Harbor rules replace true birth, death and event dates: events land inside a 2100-2200 window, and each patient carries an anchor_year plus anchor_age recorded at that year. Seasonality, aging across visits and spacing between admissions all stay computable - you reason in anchor years rather than calendar years.
Does the dataset include clinical notes and images?
Yes. A dedicated notes module carries discharge summaries and radiology reports as free text with supporting detail tables, the imaging companion pairs 377,110 chest X-ray images across 227,835 studies with their reports, and an ECG module links waveform records to machine measurements. Text, image and signal sit alongside the structured record instead of apart from it.
Can scheduled deliveries be turned into a longitudinal panel?
That is the intended shape. Every cut reuses identical field names and the same subject-to-stay identifier spine, so successive snapshots concatenate into time-ordered panels without column remapping. Teams typically anchor on the core record and layer notes or imaging cuts as separate frames keyed to the same subjects.
Datasets that pair with this one
- MIT Laboratory for Computational Physiology / PhysioNet - source profile The lab behind the corpus - everything else its data operations publish, profiled end to end.
- PhysioNet - Research Resource for Complex Physiologic Signals Several hundred curated waveform, ECG and signal collections - breadth to set beside one record's depth, compared head to head in our [side-by-side](/compare/mimic-iv-medical-information-mart-for-intensive-care-vs-physionet-research-resource-for-complex-physiologic-signals).
- The Cancer Imaging Archive (TCIA) Oncology imaging at archive scale - the counterpart shelf when the modality is tumor, not telemetry.
- ONC Health IT Dashboard and Open Data National health-IT adoption and certification statistics - the population-level view beside one hospital's record-level truth.
- Best health care technology datasets Where this record ranks among the industry's six cataloged sources - it tops the list at 10/10.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.