Datadory notebook

MAUDE adverse event database: where 25.7 million device reports become usable rows

Datadory delivers health care equipment data covering the largest device-safety record in medicine: the MAUDE adverse event database - 25,711,469 reports of device malfunctions, injuries and deaths reaching back to roughly 1992, one row per report with nested device, patient and narrative detail - delivered daily, weekly, or hourly, your call.

1,744 datasets. Pick your catch.

What is the MAUDE adverse event database?

When a medical device may have contributed to a death, a serious injury or a malfunction, someone files a medical device report. Manufacturers, importers and device user facilities file because they must; health professionals, patients and consumers file because they saw something. The releasable results accumulate in MAUDE - the Manufacturer and User Facility Device Experience collection - and assembled into rows it is the largest single body of device-harm evidence medicine has: 25,711,469 reports at the August 2026 research pass, growing by several hundred thousand more every year, with publicly releasable records reaching back to roughly 1992.

Datadory delivers it as the openFDA Device Adverse Events (MAUDE) API dataset, scored 10 out of 10 on our rubric - one of only 145 perfect scores among the 1,744 datasets we catalog, against a catalog-wide average of 7.81. In the six-record Health Care Equipment pool, nothing else competes on volume: the [manufacturing census]( /datasets/health-care-equipment/openfda-device-registration-listing-api) alongside it holds 333,804 establishment-product pairs, and the global policy panel adds roughly 4,700 country-year cells.

Scale is half the story. Structure is the other half, and it is why the corpus behaves like a product rather than an archive. Every report lands as one row keyed on mdr_report_key, typed by outcome in event_type (Death, Injury, Malfunction, Other or no answer), stamped twice - when the event happened and when regulators received it - attributed through report_source_code, and carrying one nested array per device involved, one per patient affected, plus the filer's own narrative. On top of the raw submissions sits a harmonized block joining each suspect device to its regulated name, risk class and regulation number, which is what turns twenty-five million accounts into something a query can group.

What does a delivered MAUDE report look like?

One real record, captured during the August 2026 research pass and flattened for reading - the oldest publicly releasable shape in the corpus:

# mdr_report_key 10 - the oldest publicly releasable report shape
mdr_report_key : 10
event_type     : Injury
date_of_event  : 19920220
date_received  : 19920310

device.brand_name        : N/A
device.generic_name      : MANUAL HOSPITAL BED
device.model_number      : 720
device.device_report_product_code : FNJ

openfda.device_name       : Bed, Manual
openfda.device_class      : 1
openfda.regulation_number : 880.5120

Three things are visible in eleven lines, and each one matters downstream. First, the two-clock structure: the event happened 1992-02-20 and the report reached regulators 1992-03-10 - an eighteen-day lag, and the gap between date_received and date_of_event measures exactly that across the whole corpus. Second, the device resolves twice: once as filed (MANUAL HOSPITAL BED, model 720, product code FNJ) and once harmonized (Bed, Manual, class 1, regulation 880.5120). That pairing is why product-code joins hold - the reporter's spelling varies, the code does not. Third, brand_name reading N/A is not missing data to impute; it is how reprocessed single-use devices are legitimately reported. Multiply that distinction across 25.7 million rows and vary the product code, and you have post-market surveillance.

What fields does each MAUDE report include?

The delivered dictionary runs eighteen documented fields, grouped four ways, with definitions verified against the agency's own field reference rather than inferred from column headers.

Report identity and outcome. mdr_report_key is the unique identifier - an 8-digit string whose final digit is a checksum, which catches transcription errors before they enter a pipeline. event_type sorts every report into Death, Injury, Malfunction, Other or no answer, the single most-used split on the corpus. date_of_event and date_received arrive as YYYYMMDD strings, and report_source_code separates mandatory manufacturer reports from user-facility, distributor and voluntary submissions - the weighting variable nearly every cut needs.

Device detail. The nested device[] array carries brand_name, generic_name, model_number, catalog and lot numbers, an implant flag, the redacted public UDI (present only on reports from October 2022 onward), and the three-letter device_report_product_code that classes the device under 21 CFR Parts 862-892.

Patient detail and narrative. patient[] records age, sex, weight and the reported problems and outcomes, while mdr_text[] preserves the filer's narrative - the layer structured fields cannot replace, because it explains what actually happened rather than which box got ticked.

Harmonized classification. The enrichment block resolves each device to its regulated name, risk-based class and regulation number such as 880.5120, so bare product codes become analyzable categories without a lookup project.

One design fact to plan around: fields switch on mid-history. date_of_event exists only from 2006 onward and the public UDI from October 2022, so longitudinal designs tolerate columns that start partway through the timeline - the delivered tables preserve that absence rather than zero-filling it away, so missing stays distinguishable from zero.

What can you build with the MAUDE database?

The working questions are surveillance questions, and each maps onto named fields rather than prose guessing:

  • Signal detection by product code. Group reports by device_report_product_code over time and watch acceleration - a category's malfunction mix shifting shows up as a count long before it shows up in an announcement.
  • Malfunction-versus-harm triage. event_type separates hardware failure from patient harm in one field, which is why death-to-malfunction ratios read straight off this corpus rather than any other.
  • Manufacturer surveillance. Pull one firm's history through manufacturer_name, resolved against the registration census's FEI-numbered establishments so subsidiaries and dba variants roll up correctly, and line it against recall actions on the same codes.
  • Risk modeling and narrative NLP. Millions of outcome-labeled rows carrying device attributes, patient details and long-form narratives make a training corpus few domains match for volume - with the reporting biases in the next section priced in rather than discovered late.
  • Litigation and investment early warning. Accelerating report counts per manufacturer precede recall notices and complaints; desks watching event_type mixes by firm see the curve before the filing.

Two habits prevent rework regardless of the job. Keep the nested arrays intact until the last transformation - flattening early destroys the product-code and manufacturer detail every downstream group-by depends on. And deduplicate before counting events, because the same incident can generate several reports from different reporters.

How deep and wide does MAUDE coverage run?

Geography - US reports plus foreign reports involving devices marketed in the United States, with manufacturer and distributor addresses recorded worldwide. It is the widest device-safety lens any single country publishes, but it is not a global census: a foreign incident enters only when a US-marketed device is involved.

Temporal depth - publicly releasable reports run from roughly 1992 to the present even though the documentation prints a later window, so decade-scale trend lines are real rather than reconstructed. Several hundred thousand new reports arrive each year, which means the difference between two consecutive deliveries is itself a monitoring signal - new reporting patterns become observable as history accumulates.

Grain - one record per adverse event report, nested per device and per patient involved, plus per-report narrative text. Reports are neither deduplicated nor validated cases: mandatory and voluntary reporters can file separately on the same incident, which is a feature for volume analysis and a trap for incidence math.

Set against the rest of the pool the contrast is stark. The registration census documents its establishment map from 2007 onward, and the WHO policy panel ties its observed values to survey waves ending in 2021 - at event grain and at this depth, the harm ledger has no rival in the slice.

What can MAUDE report counts not tell you?

Four limits deserve internalizing early, because they separate defensible claims from dashboards that quietly overreach. All four come straight from the publisher's own framing, and none is smoothed over in delivery.

  • Reports are submissions, not verified cases. Causality cannot be established from a report - the device may be implicated, suspected or merely present, and the agency states plainly not to rely on the corpus for medical decisions.
  • Under-reporting makes rates unknowable. Known under-reporting means incidence and prevalence cannot be calculated from these figures. Counts measure signal volume, never population rates.
  • Reporters self-select. Mandatory filings dominate the stream while voluntary submissions skew toward severe outcomes, so report_source_code belongs in nearly every cut.
  • Duplicates exist by construction. The same incident can produce several reports from different filers, so deduplicate on the report key before counting events, and corroborate any harm estimate with independent utilization data such as Medicare utilization and payment summaries.

Treat counts by product code, manufacturer and event type as directional signals and the corpus rewards you; treat them as epidemiology and it will embarrass you in review.

Which datasets complete the device-safety picture?

Who builds on the MAUDE database?

  • Data scientists and ML engineers train device-event risk models on outcome-labeled rows with device, patient and narrative detail attached - the corpus rates relevance 3 of 3 for the persona, and the workflows live at data scientists use cases.
  • Developers and data-product builders power recall-alert and safety features from a corpus deep enough that coverage gaps rarely force a second source.
  • Journalists, academics and students investigate device injuries with citation-grade sourcing - every figure traces to the agency itself through a stable report key, the pattern documented at journalists academics use cases.
  • Competitive intelligence and product teams watch rivals' adverse-event profiles inside shared product codes, before the same patterns surface in recall notices or dockets.
  • Quality and regulatory teams benchmark internal complaint intake against the sector-wide stream to see whether their own volumes track category norms.
  • Investors and quants read accelerating report counts per manufacturer as an early input to recall and litigation screens.

Whatever the desk, the habit is identical: work the coded fields first, read the narratives wherever the counts surprise you, and never present a count as a rate.

How is the MAUDE database delivered through Datadory?

API, files, or your warehouse. Daily, weekly, or hourly.

Name the product codes, manufacturers and date ranges you care about when you request a sample; the ongoing arrangement follows once the rows validate against the documented schema.

Rows arrive typed - dates as dates, enums as ordered categories, nested arrays flattened into shapes a warehouse loads without a parsing project, or kept nested if your tooling prefers it. Successive deliveries are kept rather than overwritten, which converts a growing submission stream into diffable history: the reports that landed between two loads surface as events instead of vanishing into the running total.

Where should you start?

Start with the anchor record, the openFDA Device Adverse Events (MAUDE) API dataset page - full field dictionary, verified sample rows and coverage chips included - and request a sample cut to your product codes before anything ongoing is committed. The vocabulary decodes further on the MAUDE glossary entry and the MAUDE adverse event reporting entry.

Match the device-safety question to the record and the field
Question you are askingRecord to pullField to reach for
Is this device category getting safer or worse?openFDA Device Adverse Events (MAUDE) APIevent_type grouped by device_report_product_code over time
Which firms dominate harm reports in my code?openFDA Device Adverse Events (MAUDE) API + Registration & Listingmanufacturer_name resolved through FEI-numbered establishments
What did the agency do after the reports?openFDA Device Recalls APIreason_for_recall and classification joined on firm and product code
When did this model reach the market?openFDA Device 510(k) Clearances APIdecision date and predicate device via the K number
Which exact model version was involved?openFDA Unique Device Identifier (GUDID) APIbrand, model and catalog fields keyed on the device identifier
Who pays for the care afterward?CMS Data Hub - Medicare & Medicaid Datasetsutilization and payment summaries joined at the analysis layer

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Health Care Equipment United States plus foreign reports involving US-marketed devices

openFDA Device Adverse Events (MAUDE) API

Health Care Equipment Worldwide - US and foreign establishments registering to…

openFDA Device Registration & Listing API Data

Health Care Equipment Up to ~143 WHO Member States per density indicator with…

WHO Medical Devices – Definitions and Policy Repository

IndicatorCode · IndicatorName · SpatialDim …+4 more

Health Care Supplies United States enforcement actions against FDA-registered…

openFDA Device Recalls API

Health Care Supplies Devices marketed in the United States

openFDA Device 510(k) Clearances API

k_number · applicant · contact …+9 more

Health Care Supplies Devices distributed in the United States

openFDA Unique Device Identifier (GUDID) API

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

How large is the MAUDE adverse event database?

25,711,469 reports at the August 2026 research pass, growing by several hundred thousand medical device reports a year, with publicly releasable records running from roughly 1992 to the present. It is the largest device-safety corpus in Datadory's 1,744-record catalog and one of only 145 datasets to score a perfect 10.

What fields does each MAUDE adverse event report carry?

One row per report: an identifier whose final digit is a checksum, an event_type of Death, Injury, Malfunction, Other or no answer, paired event and receipt dates, the reporting source, a nested array per device involved, a nested array per patient affected, the filer's narrative, and a harmonized block adding regulation number, risk class and specialty.

Can MAUDE report counts be turned into injury rates?

No. Reports are submitted, not validated; causality cannot be established from a single report, and known under-reporting prevents calculating incidence or prevalence. Treat counts by product code, manufacturer and event type as directional signals, deduplicate multi-reporter incidents, and corroborate harm estimates with independent utilization data.

Does the MAUDE database connect adverse events to recalls?

Not on its own - a report carries no recall-status field. Pair the Device Recalls record, 58,999 classified events, joining on recalling firm and establishment identifier with the shared three-letter product code, to attach reasons, root causes and hazard classifications to individual reports.

Who uses the MAUDE database day to day?

Post-market surveillance teams isolating malfunctions versus injuries by product code, data scientists training device-risk models on millions of labeled rows, journalists tracing device injuries to citable sources, and competitive-intelligence and investor desks reading report acceleration ahead of recall and litigation news.

How is the MAUDE database delivered through Datadory?

As typed rows - API, files, or straight into your warehouse, daily, weekly, or hourly, your call. Nested arrays flatten into warehouse-ready shapes, successive deliveries are kept so new reports arrive as diffable history, and samples come pre-cut to the product codes, firms and date ranges you name.