Datadory notebook
Is openFDA Device Data Free for Commercial Use?
1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.
1,744 datasets. Pick your catch.
What does each of the four device corpora contain?
Four tables, one join plan - market entry, market exit, product identity and post-market evidence.
Device 510(k) Clearances is the market-entry ledger (quality 10). One record per premarket notification - 175,814 of them from the 2026-08-10 load - carrying applicant identity and address, device name, three-letter product code, advisory-committee panel, receipt and decision dates, decision code and predicate tracking back to 1976.
Device Recalls is the enforcement ledger (quality 10). One record per recall event among 58,999 classified since November 1, 2002: recalling firm with FEI number, free-text product description, reason for recall, root cause, corrective action taken, affected quantity, distribution pattern down to named states, plus k_number and pma_number arrays tying each event back to its premarket submissions.
Unique Device Identifier (GUDID) is the master-data spine (quality 9). One record per marketed model version among 5,083,948 filed since the UDI rule's 2013 phase-in: brand and catalog numbers, version/model, labeler company with DUNS, primary and previous DIs, GMDN terms with plain-language definitions and implantable flags, FDA product codes enriched with device class and regulation number, sterilization attributes, distribution status and around twenty labeling booleans.
Device Adverse Events (MAUDE) is the evidence corpus (quality 10). One record per report among 25,711,469 spanning roughly 1992 forward, several hundred thousand new reports landing each year: event_type sorting Death, Injury, Malfunction and Other into one field, nested device arrays with brand, model, lot and UDI-DI, patient demographics and outcomes, narrative text, and a harmonized block joining every device to its class and regulation number.
What do delivered rows look like?
Rows exactly as they land - one corpus per block, keys intact:
# 510(k) clearance row
k_number : K142820
applicant : Abb Optical Group, LLC
device_name : BIOLENS Sphere (mangofilcon A) Soft (hydrophilic)
Contact Lens for Daily Wear [and variants]
product_code : LPL panel : Ophthalmic
date_received : 2015-04-30 decision : SESE (Substantially Equivalent)
# recall event row
res_event_number : 86352
recalling_firm : Ansell Healthcare Products LLC
product : MICROFLEX Diamond Grip Examination Gloves, MF-300
reason : shipped inadvertently without testing to verify barrier integrity
status : Terminated initiated : 2020-08-19 quantity : 1312 Cases
# GUDID device master row
brand_name : MICROFLEX
version_or_model : XC-310-XL
company_name : Ansell Healthcare Product
primary_di : 00769799310140
commercial_distribution_status : In Commercial Distribution
# MAUDE adverse event row
mdr_report_key : 10
event_type : Injury
date_of_event : 19920220 date_received : 19920310
device.generic_name : MANUAL HOSPITAL BED product_code : FNJ
openfda.device_class : 1 regulation_number : 880.5120Read the four together and a device-intelligence product writes itself: the glove brand in the master table is the same firm recalling 1,312 cases in 2020, reachable through k_number arrays; the contact lens cleared through the ophthalmic panel sits beside every other lens ever cleared through it; the 1992 bed report shows the two-clock structure - event date versus received date - that reporting-lag metrics measure across millions of rows. Request a sample naming your product codes or brands and the same four shapes come back filled for your scope.
Where does coverage run, and at what grain?
Geography - United States throughout, which runs deeper than it sounds: foreign manufacturers selling into the US market file too, so the device master doubles as a directory of offshore production serving American healthcare, keyed by company name and DUNS. Recall rows add distribution_pattern, naming the states and territories a recalled lot actually reached.
Temporal - five decades on the entry side: 510(k) notifications from 1976, recalls classified since November 1, 2002 with firm-initiated corrections added from January 3, 2017, device identifier filings from the 2013 UDI phase-in, adverse-event reports from roughly 1992. Each corpus lands on its own schedule and every delivery carries its load date, because any stored count is a floor - filings arrive continuously.
Granularity - one row per regulatory object: one notification, one recall event, one model version, one report. No sampling layer, no aggregation imposed upstream, so counts recomputed from delivered rows reconcile against published totals exactly.
Two seams deserve naming before they surprise anyone. Once FDA classifies a recall, the record is not revised beyond Enforcement Report corrections, so termination dates are authoritative even where late detail never arrives - structure to model, not breakage to fix. And the adverse-event stream mixes mandatory and voluntary reporting: absence of reports is not absence of risk, and the corpus cannot support incidence or prevalence arithmetic. Both caveats travel inside the delivery as documented fields rather than footnotes someone finds later.
What can you ship commercially once the device shelf is on site?
Four workflows pay for themselves fastest.
Procurement catalog builds. Join internal SKUs onto standardized brand, model, GMDN term, product-code, sterilization and implantable attributes on the primary DI - so purchasing stops stocking three spellings of one examination glove, and delisted models stay visible through commercial_distribution_end_date instead of vanishing quietly.
Competitive market-entry monitoring. Diff successive clearance deliveries for new applicants and product codes across 175,814 records reaching back to 1976; new 510(k)s surface months before press releases do, with the predicate chain showing exactly whose market share a newcomer claims to be equivalent to.
Supplier risk scoring. Screen counterparties against the 58,999-event recall history by recalling firm, classification, root cause and distribution pattern before contracts are signed, then keep the feed current so newly recorded events diff cleanly into yesterday's rows.
Model training and post-market surveillance. Millions of outcome-labeled adverse-event rows carrying device, patient and narrative text pair with recall classifications to flag devices whose safety signals precede corrective action - training volume few domains match.
One boundary survives any delivery arrangement: permissive rights are legal, not clinical. The records describe what was filed, not validated safety judgments, so products quote them with provenance attached rather than presenting them as medical guidance.
Which datasets complete the picture beside the four FDA corpora?
Clearances, recalls, identifiers and adverse events cover the regulated life; five neighbours fill the edges they leave open.
- ClinicalTrials.gov API v2 (quality 10) - the forward-looking half: roughly 600,000 registered studies worldwide with 89,921 matching a device-intervention query, so the evidence pipeline competitors are building shows up years before its products do.
- Device Registration & Listing API (quality 9) - the facility map: 333,804 records linking establishments - contract manufacturers, sterilizers, repackagers, importers - to the devices they list for US sale, with FEI numbers and full addresses for supply-chain mapping.
- U.S. Census Bureau Data Portal (quality 8) - the market-sizing layer: County Business Patterns establishment counts since 1964 and five-yearly Economic Census deep cuts, one observation per geography per NAICS industry per year.
- World Bank Health Topic DataBank (quality 9) - the demand context: 658 health indicators across 295 economies, hospital beds and health spending included, annual country-year observations joining to trade tables on ISO3 codes.
- FDA Medical Devices Portal (quality 6) - the canonical index of roughly two dozen enumerated device databases, useful when a workflow needs a corner of the regulation none of the structured corpora exposes yet.
Who builds on FDA device data?
Ranked by how directly the shelf answers the day job.
- Procurement and supply-chain teams build item masters off 5,083,948 device master records and screen suppliers against the 58,999-event recall history - catalog hygiene and counterparty screening from the same join key. See developers & builders use cases.
- Competitive-intelligence product teams run the clearance ledger as market-entry alerting, decomposed by applicant, product code and predicate chain across five decades of decisions. See competitive intel product teams use cases.
- Data scientists and ML engineers get labeled volume: 25.7 million outcome-tagged adverse-event reports joined to device classes through stable product-code keys, plus trial outcomes for forward labels. See data scientists use cases.
- Market researchers and consultants size categories by distinct models per GMDN term - assortment density computed from a complete denominator rather than a sample. See market researchers use cases.
- Investors and analysts read clearance velocity and recall classifications as leading indicators on medtech names, with five decades of entry history giving any backtest its regime variety. See investors & quants use cases.
Why get FDA device data through Datadory?
Because the hard part was never the first extract - it is the tenth. Dates arrive as day-first strings in one corpus and eight-digit integers in another. Nested device, patient, nomenclature and submission arrays need flattening onto stable keys before any warehouse loads them. Classification arrives twice - as filed and harmonized - and pipelines have to choose. Recent adverse-event windows keep gaining late-arriving reports after they close, so a snapshot quietly misleads anyone building trend lines. Each is survivable once; none is fun to re-solve in every notebook.
Datadory normalizes before delivery: real dates throughout, nested blocks flattened into columns while the original nesting travels alongside, both classification layers kept rather than derived, and append-only captures stacked so new model versions, recall terminations and late adverse events arrive as diffable history instead of overwrites.
Where to go next
This page covers one question inside the 20-dataset Health Care Supplies pool - eight primaries plus twelve cross-listed neighbours. Start with the pillar Health Care Supplies data guide for the full tour with field dictionaries, sample rows and quality scores side by side, or take the scored short version in the best health care supplies datasets.
When you want real rows instead of descriptions, request a sample - the field dictionary travels with it.
| Dimension | Coverage |
|---|---|
| Geographic | United States throughout; foreign labelers and applicants filing for the US market included, so the device master doubles as an offshore-production directory; recall rows carry distribution_pattern naming states reached |
| Temporal | 510(k) clearances 1976-present; recall events classified November 1, 2002-present with corrections added from January 3, 2017; device identifier filings 2013-present; adverse-event reports roughly 1992-present |
| Granularity | One row per regulatory object - notification, recall event, model version, report - with no sampling or aggregation layer, so recomputed counts reconcile against published totals |
| Record | Layer | Coverage | Commercial fit |
|---|---|---|---|
| U.S. Census Bureau Data Portal | Market sizing (quality 8) | Establishment counts since 1964; Economic Census five-yearly 1997-2022 | Category sizing by NAICS industry across geography tiers |
| World Bank Health Topic DataBank | Demand context (quality 9) | 658 indicators across 295 economies, annual country-year series | Healthcare-demand context for international market models |
| Hugging Face Datasets Hub | Machine-learning layer (quality 7) | 1,012,419 public dataset repositories, terms set per repository | Clinical-text corpora with per-repository terms to verify before shipping |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
openFDA Device 510(k) Clearances API
k_number · applicant · contact …+9 more
openFDA Device Recalls API
openFDA Unique Device Identifier (GUDID) API
openFDA Device Adverse Events (MAUDE) API
ClinicalTrials.gov API v2 - Studies Database
openFDA Device Registration & Listing API Data
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Which device corpora should a commercial build start with?
Follow the decision being fed. Procurement catalogs start from the 5,083,948-record device master; market-entry monitoring starts from the 175,814-record 510(k) ledger; supplier-risk screens start from 58,999 recall events; surveillance and model training start from 25.7 million adverse-event reports. Most products end up needing all four, because the DI and product-code keys connect them.
Can I use FDA device data to train machine-learning models?
Yes on the regulatory corpora: millions of outcome-labeled adverse-event reports, five decades of clearance decisions and recall classifications arrive with no AI-specific restriction, and ClinicalTrials.gov adds 89,921 device-intervention studies under the same publishing logic. One embedded asset carries separate terms - the GMDN nomenclature surfaced on device records cannot be extracted to build competing terminology services, though using it inside your own joins and labels is fine.
Can a delivery be scoped to specific devices or firms?
Yes. Name the brands, GMDN terms, FDA product codes, recalling or applicant firms and date windows when you request the sample, and it ships filtered to that scope in the documented field shape - a single product-code slice or the full four-corpus estate. The ongoing feed follows the same structure, so whatever is prototyped on the sample survives into production unchanged.