openFDA Drugs@FDA API

Datadory delivers openFDA Drugs@FDA API data covering the approval biography of the American medicine cabinet: 29,273 applications reaching back to 1939, each nesting product rows - brand name, ingredients, dosage form, route, marketing status, therapeutic equivalence codes - inside submission rows that timestamp every original approval and supplement. Delivered as API, files, or a warehouse load - daily, weekly, or hourly, your call.

What is the openFDA Drugs@FDA API?

Every product on an American drugstore shelf has a paper trail that starts before the product exists, and this is where that trail lives. The openFDA Drugs@FDA API is the agency's official approval record made addressable: 29,273 applications at the August 2026 count, covering most human drug products approved since 1939 - prescription brand names, the generic challengers that followed them, many therapeutic biologicals, and over-the-counter products alike. One record sits over one application, with product rows and submission rows nested inside it.

Its position in the openFDA drug family decides what it is for. The openFDA NDC Directory API lists what is marketed (137,206 listings); the Drug Label corpus carries what packages claim about themselves (261,996 SPL documents); FAERS logs what went wrong afterward (20,692,690 reports). This is the only feed among the five that holds the approval event itself - who applied, what they filed, when the agency said yes, and whether a copy may substitute at the counter. Source: U.S. Food and Drug Administration (openFDA). Get a sample of this dataset and pull the approval biographies behind your own categories.

What does a sample approval history look like?

One real application, flattened for reading - an original approval followed by a labeling supplement, both captured during the August 2026 research pass:

application_number       N076422
sponsor_name             (applicant firm)

submissions[0]           ORIG · submission_number 1
                         submission_status        AP (approved)
                         submission_status_date   20040806
application_docs[0]      id 8910 · type Letter · date 20040806

submissions[1]           SUPPL · submission_number 4
                         submission_class_code_description   Labeling
                         submission_status        AP (approved)
                         submission_status_date   20060111

Read it as a clock instead of paperwork. The original reached approved status on 20040806 with its approval letter riding along as document 8910, typed Letter; a fourth-round labeling supplement cleared on 20060111 - roughly seventeen months between the first yes and the labeling revision. That interval is the unit of analysis here: subtract status dates inside one application and you have approval lag; stack the applications sharing a molecule and you have a generic-entry curve. Two honesty notes before anyone builds on it - firm-name strings were truncated in the captured payload, so confirm spellings before exact-match joins, and linked documents attach mostly for products approved since 1998. Your sample confirms current totals and exact header spellings before any commitment.

What fields does the dataset include?

Fifteen columns carry the working payload, all verified against live records - the table below gives each with its type, definition and a genuine example:

Additional fields on request. Product rows repeat once per marketed presentation inside an application, and submission rows repeat once per regulatory filing, so a single application fans out into a small family of child tables - the nesting is the design decision to plan around, and deliveries arrive pre-flattened if you would rather see one wide row per product or per submission. Because the therapeutic equivalence column is the one everyone asks about, ask the sample to pin down how completely it populates across your molecule set rather than assuming it fills evenly. Get a sample.

Where does coverage reach?

  • Geography: the United States regulatory perimeter. These are the approvals that gate American marketing, keyed to applicant firms rather than storefronts - there is no state or store dimension beneath the national level.
  • Temporal: approvals reaching back to 1939, with labels, approval letters, reviews and patient information mostly attached for products approved since 1998. Each submission row states its own status date, so any delivery window you keep becomes a longitudinal archive rather than a snapshot.
  • Granularity: one record per application, exploding downward - product rows repeat per marketed presentation, submission rows repeat per filing, and documents hang off individual submissions.

That combination makes this the regulatory-biography spine of the drug retail data hub: catalogs describe assortment, prices describe demand, recalls describe exits - and this feed describes the permission each product needed in the first place.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the slice, pick the cadence, pick the landing zone. Because an application keeps accumulating submissions across a product's whole life, keeping consecutive deliveries is what converts a reference lookup into a research corpus: new supplements arriving, tentative approvals converting, marketing statuses flipping. Diff two deliveries and the week's new filings fall out on their own - no re-pulls, no missed milestones. The sample comes first, so record shapes and totals are settled facts before any commitment. Get a sample of this dataset and wire it into your own stack.

Who builds on approval history, and for what?

  • Generic-entry timing and pricing-pressure modeling - therapeutic equivalence ratings plus submission status dates turn the corpus into entry windows per molecule, ready to lay under a price panel; methods continue on our ML model training page.
  • Event studies around approvals - reconstruct the filing chain from the original application to the latest supplement and align it against sales or volume series; playbooks sit on the data scientists use cases page.
  • Competitive watch on challenger entrants - a new ANDA application naming an incumbent's molecule shows up as rows long before it shows up as headlines; screens live on the competitive intel use cases page.
  • Regulatory-standing checks before listing - marketing_status separates Prescription, OTC, Discontinued and None (tentative), which is the difference between stocking a product and stocking a placeholder.
  • Citation-grade verification - writers and researchers check approval dates and sponsors against the official record with the agency's own letters attached; reporting frames sit on the journalists academics use cases page.
  • Category milestone marking - analysts stamp approval milestones into category reviews so competitive resets are dated rather than remembered.

Which personas get the most value?

Data scientists and quant researchers get the rare regulator-grade corpus whose nesting is a feature rather than a chore - flatten submissions into one timeline per application and the event study falls out; catalyst-timing plays collect on the investors quants use cases page. Developers and builders get flat JSON records that load without parsing gymnastics, with coded values - AP, TA, ORIG, SUPPL - that resolve through short dictionaries instead of bespoke mappings; integration notes sit on the developers builders use cases page. Market researchers and consultants get dated category resets for briefings that survive scrutiny; patterns sit on the market researchers use cases page.

Which datasets and notes pair with it?

Patent-and-exclusivity boundary worth internalizing early - patent and exclusivity details live in the agency's separate Orange Book data rather than in Drugs@FDA, so full patent-cliff work pairs the two instead of expecting columns this feed never carried. The equivalence code speaks to substitutability at the counter, not to patent expiry.

Where to go next - the rail below collects the identity spine, the label and safety counterparts, the removal ledger, the consumer-price layer, the comparison against alternatives, and the glossary entries decoding application vocabulary, starting with Drugs@FDA approval application and the National Drug Code explained.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary for the openFDA Drugs@FDA API - fifteen verified columns; examples are genuine values from a live application record
fieldtypedefinitionexample
application_numberstringFDA application identifier, prefixed NDA, ANDA, BLA or BN depending on application type - the scoping key for every join.N076422
sponsor_namestringName of the applicant firm holding the application; the supplier and competitor watchlist column.(applicant firm)
products.product_numberstringProduct ordinal within the application; one application nests many marketed presentations.-
products.brand_namestringMarketing name of the product when branded; blank for unbranded generics.-
products.active_ingredientsstringIngredient name plus strength pairs for the product; the molecule-level grouping column.-
products.dosage_formstringDosage form of the marketed product.TABLET
products.routestringRoute of administration.ORAL
products.marketing_statusstringCurrent standing of the product: Prescription, OTC, Discontinued, or None (tentative approval).Prescription
products.reference_drugstringWhether the product serves as the reference against which equivalents are judged.-
products.te_codestringTherapeutic Equivalence rating indicating substitutability at the counter - the generic-entry column.-
submissions.submission_typestringType of regulatory filing: ORIG for the original application, SUPPL for supplements.ORIG
submissions.submission_statusstringReview outcome of the filing: AP (approved) or TA (tentative approval).AP
submissions.submission_status_datedateDate the filing reached its status (YYYYMMDD); the column approval timelines are built from.20040806
submissions.submission_class_code_descriptionstringClass of the submission, such as Labeling, Review or Risk REMS.Labeling
submissions.application_docsarrayLinked documents with id, date and type (Letter or Review), attached mostly for post-1998 approvals.doc 8910 · Letter

Coverage at a glance

DimensionCoverage
GeographyUnited States regulatory perimeter - the approvals that gate American marketing, keyed to applicant firms rather than storefronts
TemporalApprovals since 1939; linked letters, reviews and patient information mostly since 1998; every submission row carries its own status date
GranularityOne record per application, with product rows repeating per marketed presentation and submission rows repeating per filing
Scale29,273 applications counted August 2026

Questions buyers ask

How many drug applications does the openFDA Drugs@FDA dataset cover?

29,273 applications at the August 2026 count, covering most human drug products approved since 1939 - prescription brands, generics, many therapeutic biologicals and OTC products. One record holds one application with product and submission rows nested inside, so the flattened row count fans out well past the application count.

What does a therapeutic equivalence (TE) code mean in these rows?

It grades whether one product is therapeutically equivalent to its reference - the agency's own substitutability judgment carried straight onto the product row, and the difference between a generic pharmacies treat as interchangeable and one that is not. The code lives on product rows rather than application rows, so flattening decisions decide whether you can read it at all.

How do submission rows build an approval timeline?

Every filing arrives as its own submission row - ORIG for the original application, SUPPL for supplements - carrying a status of AP (approved) or TA (tentative), a class description such as Labeling, and the YYYYMMDD date the filing reached that status. Sort the rows inside one application and the product's regulatory biography reads chronologically, from first yes to latest labeling change.

What does a None (tentative) marketing status indicate?

That the product holds a tentative approval rather than a final one - reviewed and cleared on its merits but not yet in marketing, which is why the status column distinguishes it from plain Prescription and OTC standings. For entry-timing work these rows matter as much as approved ones: they mark challengers already through the review process, waiting on something else.

Does this dataset include patents and exclusivity dates?

No - patent and exclusivity details sit in the agency's separate Orange Book data, so full patent-cliff analysis pairs the two sources rather than expecting columns this feed never carried. Drugs@FDA holds the application chain, products and submissions; the Orange Book holds the intellectual-property clock. The join runs on product identity, and combining them is standard practice for entry-window modeling.

How do approval rows join to other drug retail datasets?

Through product identity. Brand names and active ingredients resolve to National Drug Codes against the NDC Directory, attaching approval biographies to cash-price panels, shelf assortments, adverse-event profiles and recall ledgers keyed to the same products. Flatten products and submissions first - one row per product, one per submission - or the join duplicates silently.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing