KEGG DRUG Database
Datadory delivers KEGG DRUG pharmaceuticals data covering 12,896 approved drugs across Japan, the USA and Europe - each a flat-file record keyed by D number with names, formulas and masses, pharmacological efficacy, therapeutic targets, metabolizing enzymes, transporters, ATC and national classification codes, label links and interaction flags, delivered daily, weekly, or hourly.
What is the KEGG DRUG Database?
One flat-file entry per approved drug, unified by molecule rather than by paperwork. The KEGG DRUG Database held 12,896 entries as of 20 August 2026 - every drug approved in Japan, the USA or Europe, each identified by a stable D number and unified by chemical structure or active-ingredient component, so combination products and their ingredients resolve to connected records instead of unrelated strings.
The annotation depth is what separates it from approval registries: 7,335 entries carry therapeutic target annotations, 1,395 list drug-metabolising enzymes, and 274 list transporters, while label linkage reaches 2,550 entries into Japanese labels via YJ/YK codes and 2,167 into FDA labels via NDC codes, with indicated diseases mapped to KEGG H-number disease entries. Classification codes ride in the same records - ATC (2026 edition), USP DC (2026) and Japan's JP19 therapeutic categories - alongside pharmacogenomic biomarker annotations, prodrug flags and drug name stems.
Special entry types get their own counts: roughly 750 antibody therapeutics, 84 nucleic acid therapeutics and 23 gene therapy products. A companion resource, KEGG DGROUP, groups structurally and functionally related D numbers under DG numbers across five group types - Chemical, Structure, Target, Class, Metabolism - so prodrugs and their active parents stay linked. The whole corpus integrates with KEGG pathway maps, BRITE hierarchies, the drug interaction database and KEGG MEDICUS. Datadory packages it for delivery as API, files, or straight into your warehouse; get a sample of this dataset and read real D-number entries against your own drug list.
What do sample rows look like?
Entries ship as labelled flat-file blocks, one block per D number, every field self-describing. Warfarin sodium, entry D00564:
ENTRY D00564 Drug
NAME Warfarin sodium (USP); Coumadin (TN); Jantoven (TN)
FORMULA C19H15O4. Na
EXACT_MASS 330.0868
EFFICACY Anticoagulant, Vitamin K antagonist
COMMENT Coumarin derivative
TARGET VKORC1 [HSA:79001] [KO:K05357]
METABOLISM Enzyme: CYP2C9 [HSA:1559]
REMARK ATC code: B01AA03
DBLINKS CAS: 129-06-6The identifier-list view flattens the same corpus into two columns - D00002 mapping to Nadide (JAN/USAN/INN), nicotinamide adenine dinucleotide - which is the shape most join keys take. Cross-reference operations keep the same tab-delimited discipline: dr:D00564 atc:B01AA03 ties the entry to its ATC class, and dr:D00564 pubchem:7847630 converts it to a PubChem identifier. Read a full block top to bottom and you hold identity, chemistry, pharmacology and outward keys in one parse; exact header ordering and field presence vary per entry, and your sample settles which sections arrive for the drugs you care about.
What fields does the dataset include?
Twelve fields make up the working dictionary, and every one earns its row in the table below. The design principle is that nothing hides: targets name their genes with HSA and KO references, metabolism splits enzymes from transporters, and interactions arrive pre-flagged rather than as free text.
Additional fields on request. Everything outside this dictionary is where scoping happens: MOL and KCF structural representations and structural-formula images attached to entries, DGROUP membership records linking related D numbers under DG numbers, pathway-map and BRITE-hierarchy attachments, KEGG MEDICUS context, and the indicated-disease mappings behind label links. Pin down whichever of those your build actually needs before committing - the sample confirms precisely which blocks arrive.
Where does coverage reach?
- Geo: Drugs approved in Japan, the USA and Europe - a deliberately regulatory scope, not a geographic accident; the same tri-region frame drives the classification layer, with ATC 2026, USP DC (2026) and JP19 categories all present
- Temporal: Continuously curated reference content, carrying a last-updated stamp of 20 August 2026 at time of research; entries track approvals and withdrawals as regulators make them rather than freezing at release boundaries
- Granularity: One entry per approved drug ingredient or product, keyed by D number and grouped into DG families where entries are structurally or functionally related
That footprint makes it the annotation spine of the pharmaceuticals data hub: approval registries tell you a drug was authorised, commercial sources tell you what it sells for, and this corpus tells you what the molecule does - to which target, through which enzyme, against which other drugs.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Pick the slices, pick the cadence, pick the landing zone - the same D-numbered records arrive whichever way you take them. Because the corpus is continuously curated rather than versioned in releases, retaining consecutive deliveries turns a point-in-time lookup into a longitudinal history of new approvals, added target annotations, fresh interaction warnings and withdrawn entries. The sample comes first, so section layout, qualifier formatting and entry counts are settled facts before any commitment.
Who builds on it?
- Data scientists assemble polypharmacy and mechanism models where targets, metabolizing enzymes and interaction groups need to arrive joined rather than reconstructed - the 7,335 target-annotated entries give the join immediate yield. Sector-wide patterns sit on the data scientists use cases page.
- Pharmacovigilance and clinical teams screen candidate combinations against contraindication and precaution flags before they reach a monitoring plan, using the same records that feed KEGG MEDICUS's prescription interaction checking.
- Competitive intelligence teams watch interaction records, new target annotations and DGROUP additions as early signals of combination-therapy interest inside therapeutic areas they track; the playbook lives on the competitive intel product teams use cases page.
- Developers and builders wire a verified drug dictionary with ATC, USP and JP19 classifications into pharmacy tooling and analytics products instead of maintaining homegrown code sets; integration patterns sit on the developers builders use cases page.
- Market researchers anchor therapy-area sizing to a single canonical drug vocabulary spanning three jurisdictions, so a molecule quoted by a Japanese trade name and an American generic resolves to one entity.
Which datasets and notes pair with it?
- WHO ATC/DDD Index - the official anatomical-therapeutic-chemical hierarchy with Defined Daily Doses; KEGG's derived ATC codes join straight onto it for therapy-area rollups.
- openFDA Drug APIs - adverse-event reports, structured product labels and NDC listings covering the same molecules from the post-market side, reachable through the NDC links KEGG already carries.
- ChEMBL Web Services API - tens of millions of measured bioactivity records keyed to structures you can pre-screen here first.
- RxNorm / RxNav Browser and APIs - normalized US clinical vocabulary and RXCUI crosswalks once prescribing or claims data enters the picture.
- EMA Medicines Catalogue - centrally authorised European medicines, complementing the Japanese and FDA label links with the EU regulatory view.
Two reference points frame the choice before you commit: how the publisher positions this KEGG data inside its wider resource family, and where the KEGG DRUG Database vs KEGG REST API (DDI and Link Operations) comparison lands when the question is content versus query interface. Three glossary notes sharpen the vocabulary: how ATC classification organises drugs into therapy hierarchies, what drug interaction data encodes beyond a pair of names, and where national drug codes fit alongside D numbers and CAS entries.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
ENTRY | string | D-number identifier and entry type opening each flat-file record; the primary key every downstream system anchors to. | D00564 Drug |
NAME | string | Drug name(s) qualified by source vocabularies - USP, JP19, JAN, INN or TN trade names. | Warfarin sodium (USP); Coumadin (TN) |
FORMULA | string | Molecular formula of the active component. | C19H15O4. Na |
EXACT_MASS | number | Exact monoisotopic mass of the molecule. | 330.0868 |
EFFICACY | text | Pharmacological efficacy description. | Anticoagulant, Vitamin K antagonist |
COMMENT | text | Chemical-class or other annotation notes. | Coumarin derivative |
TARGET | text | Therapeutic target with human gene (HSA), KO and variant cross-references. | VKORC1 [HSA:79001] [KO:K05357] |
METABOLISM | text | Metabolizing enzymes and transporters acting on the compound. | Enzyme: CYP2C9 [HSA:1559] |
INTERACTION | text | Drug interaction entries flagged contraindication (CI) or precaution (P). | <contraindicated D number> |
REMARK | text | Remarks including the derived ATC classification code. | ATC code: B01AA03 |
DBLINKS | text | External identifier links such as CAS registry numbers. | CAS: 129-06-6 |
PRODUCT.GENERIC | text | Generic product listings tied to FDA/NDC label data for US drugs. | <US generic product listing> |
Questions buyers ask
What is inside one KEGG DRUG record?
A labelled flat-file block per approved drug: ENTRY carrying the D number, NAME with source-vocabulary qualifiers, FORMULA and EXACT_MASS for the molecule, an EFFICACY description, COMMENT class notes, TARGET gene references, METABOLISM enzymes and transporters, INTERACTION flags marked contraindication or precaution, REMARK holding the derived ATC code, DBLINKS with CAS registry numbers and PRODUCT.GENERIC listings for US generics.
How many drugs does it cover, and how deep do annotations go?
12,896 entries as of 20 August 2026, one per approved drug ingredient or product. Depth is countable: 7,335 entries carry therapeutic target annotations, 1,395 list drug-metabolising enzymes, 274 list transporters, 2,550 link into Japanese labels through YJ/YK codes and 2,167 into FDA labels through NDC codes.
Which countries' approved drugs are included?
Japan, the United States and Europe only - that regulatory scope is the dataset's defining boundary rather than a gap. The same tri-region frame drives the classification layer: ATC codes from the 2026 edition, USP DC (2026) categories for the American vocabulary and Japan's JP19 therapeutic classes ride in the same records.
What can I join KEGG D numbers against?
Outward bridges come built in: DBLINKS fields carry CAS registry numbers, structure-level conversion reaches PubChem identifiers, REMARK holds derived ATC codes for joins onto WHO classification data, and label links run to FDA NDC codes and Japanese YJ/YK codes. Inward, companion KEGG resources attach pathway maps, BRITE hierarchies and DGROUP families of related D numbers under shared DG numbers.
How current is the data?
The corpus is continuously curated and carried a last-updated stamp of 20 August 2026 when researched on 21 August 2026 - a one-day gap between curation and observation. Datadory delivers whatever the corpus holds on your schedule, daily, weekly, or hourly, so freshness becomes a cadence you set rather than one you discover.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.