Drug Retail Data Provider: 16 Cataloged Datasets · Head-to-head

openFDA Drug Label API (Structured Product Labeling) vs openFDA NDC Directory API

Which drug retail data provider: 16 cataloged datasets data fits your job: openFDA Drug Label API, or openFDA NDC Directory API. API, files, or your warehouse. Daily, weekly, or hourly.

Drug Retail Data Provider: 16 Cataloged Datasets United States · Document versions effective from June 2009 to present

openFDA Drug Label API (Structured Product Labeling)

Drug Retail Data Provider: 16 Cataloged Datasets United States market - products listed with FDA by labelers selling in the US · Current-market snapshot whose marketing start dates reach back decades and whose listings carry certification expiry dates

openFDA NDC Directory API

Coverage, side by side

openFDA Drug Label API openFDA NDC Directory API
Geographic United States United States market - products listed with FDA by labelers selling in the US, identified through labeler identity rather than geography
Temporal Document versions effective from June 2009 to present, with some earlier records Current-market snapshot whose marketing start dates reach back decades and whose listings carry certification expiry dates
Granularity One record per SPL label version, with section-level text arrays inside each One row per marketed product, with nested entries per package and per active ingredient

What each contains

They tie on 2 attributes. Pick by fit, not by loyalty.

openFDA Drug Label API openFDA NDC Directory API
Publisher U.S. FDA openFDA U.S. Food and Drug Administration (openFDA)
Subject What America's drug products say about themselves: full Structured Product Labeling text submitted by manufacturers and distributors for prescription and OTC drugs What America's drug products are: the registry of marketed finished drug products filed under the Drug Listing Act of 1972
Record scale 261,996 label records at the August 2026 research pass 137,206 product listings in the August 2026 export
Unit of analysis One record per SPL label version, with section-level text arrays inside each One row per marketed product, with nested entries per package and per active ingredient
Geographic coverage United States United States market - products listed with FDA by labelers selling in the US, identified through labeler identity rather than geography
Temporal character Document versions effective from June 2009 to present, with some earlier records Current-market snapshot whose marketing start dates reach back decades and whose listings carry certification expiry dates
Field dictionary 21 documented fields per record, with roughly ninety searchable label sections inside each document 21 documented fields per record, definitions verified
Delivery cadence Daily, weekly, or hourly - your call Daily, weekly, or hourly - your call
Best for Claim intelligence: what a product claims to treat, how it instructs dosing, what it warns against, verbatim Shelf intelligence: what is marketed, by whom, in what form, under which regulatory pathway and code

Or take both in one feed

Yes - this is the pairing the Drug Retail slice exists for: the ledger decides what is on the shelf, the labels say what the shelf claims.

A first-pass workflow: select the population from the directory side - every marketed HUMAN OTC DRUG in a given dosage_form and pharm_class, say - then fan out to the label side through the shared codes and pull each product's current indications, dosing instructions and warnings as text. Category-level claim maps, warning-language gap analyses and competitive copy comparisons all reduce to that one hop.

Watch four seams. First, cardinality runs one-to-many: one product listing faces many label versions, so filter to the latest effective_time per set_id before joining or the counts inflate. Second, the harmonized block is not universal - some label records carry no populated openfda block, so keep a name-based fallback for the residue. Third, maker concepts differ: openfda.manufacturer_name and labeler_name will not string-match cleanly until corporate suffixes are normalized. Fourth, respect the registers - never aggregate label text counts beside listing counts as if they were the same unit; one is documents, the other is products.

Handled together, the pair answers questions neither half can: what is marketed, in what form, by whom, saying what.

Or take both in one feed.

API, files, or your warehouse. Daily, weekly, or hourly.

Fair questions

Do the two field dictionaries overlap?

On the identifier spine, almost completely. Both carry the National Drug Code (`openfda.product_ndc` on the label side, native `product_ndc` on the directory side), both tie back to SPL document identity (`set_id` versus `spl_id`), and both resolve brand name, generic name, manufacturer and administration route. Beyond that spine they hold opposite registers: roughly ninety searchable label sections of manufacturer-written prose against twenty-one typed registry fields - sentence-borne claims versus column-typed facts.

Which one covers more ground?

Different units, so the counts do not compete. Labels outnumber listings - 261,996 against 137,206 - because every revision of a product's labeling stands as its own versioned document, while a marketed product remains one listing however often it changes. Reading one product deeply takes the label side; reading the whole American shelf takes the directory.

Which record is more current?

They move differently. The label corpus accumulates document versions from June 2009 onward, so currency arrives as newer versions of the same `set_id`. The directory is a current-market ledger: one row per marketed product, with marketing start dates reaching back decades and expiration dates certifying how fresh each listing is. Datadory ships either one on a single cadence - daily, weekly, or hourly - your call.