Datadory notebook
Drug interaction database, delivered: the pairwise corpus behind every interaction checker
Datadory delivers drug interaction data covering the whole edge layer: contraindication and precaution pairs across 12,896 approved-drug entries from Japan, the USA and Europe, mechanism context from 7,335 therapeutic-target annotations and CYP metabolism lines, verbatim US label wording from roughly 262,000 structured product labels, and ATC classification parity across 14 anatomical groups - delivered as typed, joined rows daily, weekly, or hourly, your call.
1,744 datasets. Pick your catch.
What is a drug interaction database?
The question reads like a plumbing question and lands as a content question: when software needs to know whether compound A makes compound B dangerous, what actually exists to consult?
An interaction database serves assertions pairwise - compound A alters the effect, metabolism or toxicity of compound B - so a prescribing screen queries edges instead of re-reading prose labels. Counted at the August 2026 research pass, Datadory's pharmaceuticals pool holds 22 cataloged records, and exactly two carry structured interaction data natively: the KEGG REST API (DDI and Link Operations) and the KEGG DRUG Database it sits on top of.
Everything else in the pool exists to make those edges usable. RxNorm / RxNav Browser and APIs resolves whatever name string arrived into an identity both sides of an edge agree on. openFDA Drug APIs supply the American regulatory wording. The WHO ATC/DDD Index standardizes classification, and ChEMBL Web Services API tests whether a flagged pair shares a protein. One pairwise corpus, thinner than most people expect, made valuable by what rides beside it.
What does an interaction record look like?
Three complete interaction rows, captured verbatim during the August 2026 research pass - precautions against one drug entry, D00564, warfarin sodium:
dr:D00564 cpd:C00304 P unclassified
dr:D00564 cpd:C01946 P unclassified
dr:D00564 cpd:C04931 P unclassifiedThe anatomy: queried drug, counterpart compound, then a type flag where P marks a precaution and CI a contraindication, closed out by a coarse class label. That is deliberately thin - no severity score, no mechanism text, no evidence citation travels with the pair. Production builds treat the flag as a tripwire and source the why elsewhere, which is exactly what the surrounding records are for.
Cross-reference rows are equally terse, and equally useful:
dr:D00564 atc:B01AA03
dr:D00564 pubchem:7847630The first joins the drug to the global classification vocabulary; the second bridges onto PubChem chemistry identifiers without a proprietary crosswalk. Fifteen neighbor databases ride alongside the drug corpus, answering to the same seven verbs, which is what turns a lookup file into a joinable graph.
Get a sample cut to your molecules and classes - real rows come back with the complete field dictionary attached.
Which datasets cover the drug interaction question?
No single table answers a drug interaction question end to end, so the slice gets bought as a small stack. Each piece ships through Datadory normalized into typed rows with its field dictionary attached, so the pieces join instead of merely coexisting.
The KEGG REST API (DDI and Link Operations) is the edge layer itself: contraindication and precaution pairs across all 12,896 KEGG drug entries, plus cross-walks into ATC, PubChem, ChEBI, UniProt and NCBI identifiers - five outside vocabularies reachable through conversion.
The KEGG DRUG Database carries the mechanism context around each edge: therapeutic targets on 7,335 of those 12,896 entries with human gene cross-references included, metabolizing enzymes named down to the CYP line, pharmacological efficacy text and ATC codes derived from the 2026 edition.
openFDA Drug APIs contribute the American wording: roughly 262,000 structured product labels carrying about ninety sections apiece with drug_interactions among them, and 20,692,690 adverse-event reports standing alongside for signal context.
RxNorm / RxNav Browser and APIs resolves identity: 120,000+ normalized concepts keyed on RXCUIs, with brand and generic names, strengths, dose forms and NDC bridges, so a claims feed and a prescribing screen stop arguing about names.
The WHO ATC/DDD Index sets classification parity: the official 2026 Anatomical Therapeutic Chemical classification with Defined Daily Doses across 14 anatomical groups, five levels deep, while the EMA Medicines Catalogue carries ATC columns for its roughly 2,700 centrally authorised European medicines.
And ChEMBL Web Services API adds mechanistic plausibility at scale: 24,527,044 bioactivity measurements across 18,552 targets, enough to check whether a flagged pair shares a target or a metabolizing enzyme before a warning ever reaches a clinician.
The comparison table below lays the six out side by side.
What fields does a delivered interaction record carry?
A raw pair is three columns wide; the deliverable is the enriched record. Datadory ships the pair joined to its mechanism, classification and label context, so the dictionary below is what actually lands - definitions verified against the documented schema, specialist blocks folded in on request.
Two design principles matter more than any single field. Nothing hides: targets name their genes with HSA and KO references, metabolism splits enzymes from transporters, and interactions arrive pre-flagged rather than as free text. And nothing pretends: the class label is frequently unclassified, which is precisely why the enrichment columns exist.
How do you wire interaction checking into a product?
The working questions are screening questions, and each maps onto named fields rather than prose guessing:
- Normalize incoming names. Pass user-, EHR- or claims-supplied strings through RxNorm first: RXCUI resolution is what lets a brand name, a generic twin and a combination pack land on the same edge.
- Resolve marketed products. The corpus accepts US and Japanese marketed-product codes natively, so a shelf item ties to a reference entry without building a private crosswalk.
- Pull the pairs. Hand over one identifier or a list; every pair comes back flagged precaution or contraindication before anything reaches formulary review. Batches run ten identifiers at a time, which shapes how bulk jobs get sliced.
- Enrich the hits. Join each flagged pair to its parent DRUG record - TARGET, METABOLISM, EFFICACY - so a warning can say why it fired instead of only that it fired.
- Quote the American label. Under display copy, quote the
drug_interactionssection of both products' US labels verbatim rather than paraphrasing; the section-level arrays keep indication and warning language exact. - Keep the deliveries. Retained consecutive feeds turn the corpus into longitudinal history: pairs appearing, moving between precaution and contraindication, or disappearing as labels change.
How far does interaction coverage reach, and where does it stop?
Three limits deserve internalizing early, because they separate defensible products from dashboards that quietly overreach.
The corpus leans on Japanese labels. Interaction records draw on Japanese label material, and KEGG's own consumer-facing checker targets Japanese prescriptions - precaution flags sit closer to the front of the evidence than most Western tables would place them. Pairing the edges with openFDA's roughly 262,000 US label documents is what makes the combined output defensible for a US-facing product.
Two grades are not a severity scale. CI and P separate contraindication from precaution; neither carries a magnitude, and the class label is frequently unclassified. Teams that need graded severity layer it themselves - adverse-event frequency from the 20.7-million-report corpus, boxed-warning presence in US labels, or shared-target plausibility from ChEMBL are the usual inputs. One documented limit applies to the adverse-event route: a multi-drug report supports no causal link between a specific drug and a specific reaction.
Prose checkers are a different artifact. Roughly 125,000 pharmacist-reviewed pages spanning 24,000+ prescription drugs, OTC medicines and natural products - the Drugs.com Drug Information Pages record among them - serve interaction tables written for humans reading a screen. Excellent reading, poor rows; the structured edge layer exists precisely because prose does not join.
Who builds on drug interaction data?
- Clinical and e-prescribing software teams screen prescription events pairwise before they render, quoting label wording under each warning so the alert survives a pharmacist's scrutiny.
- Pharmacy and formulary teams reconcile formularies against contraindication pairs and keep an auditable trail of what was flagged, when, under which release.
- Pharmacovigilance analysts read interaction flags beside 20.7 million adverse-event reports, asking whether complaints about a pair predate the flag.
- Data scientists engineer features from interaction density, shared-target plausibility and CYP overlap - countable columns rather than wrung-out prose.
- Competitive-intel product teams map crowded precaution territory class by class using ATC alignment, spotting where combination products will hit friction.
- Investors and quants take interaction density per drug and per class as a pipeline-risk base rate: crowded territory is a signal, and it arrives countable.
Persona-level playbooks collect on the industry pages: data scientists, competitive-intel product teams and developers and builders each get their own route through the same edges.
How is drug interaction data delivered?
As rows, not a research project. Datadory normalizes the interaction corpus into typed tables - flags as ordered categories, identifiers parsed and stable, mechanism text attached to the pairs it explains - so the enrichment step is a join rather than string surgery.
API, files, or straight into your warehouse. Daily, weekly, or hourly.
Name the molecules, classes and fields you need back when you request a sample; the ongoing arrangement follows once the rows validate against the documented schema.
Every delivery pins a release stamp, which is the quiet feature: because the underlying corpora revise pairs as labels change, retaining consecutive deliveries is what converts a lookup table into a research corpus where diffing two loads surfaces the week's interaction changes on their own. An alerting desk wants hourly checks, a quarterly formulary audit wants one clean pull, and neither setting requires knowing when any particular load landed. Sample first, schedule second.
Where should you start?
Start with the anchor record, the KEGG REST API (DDI and Link Operations) dataset page - sample rows, field dictionary and coverage chips live there - then read the mechanism layer on the KEGG DRUG Database page and the American wording on openFDA Drug APIs. The drug interaction data and ATC classification glossary entries decode the vocabulary, and the KEGG source profile shows what else this publisher ships.
This page is one thread of a wider map. The pharmaceuticals data guide walks all 22 pooled records end to end, the best pharmaceuticals datasets ranking scores the leaders side by side, and the pharmaceuticals data hub indexes everything with its coverage statement. Two sibling clusters extend this page directly - drug target associations dataset covers the compound-to-protein side of the same annotations, and RxNorm NDC mapping covers the identifier-normalization layer. Request a sample from any of them - the rows prove the rest.
| Dataset | Unit of observation | Scale and coverage | What it adds |
|---|---|---|---|
| KEGG DRUG Database | One flat-file record per approved drug, keyed by D number | 12,896 approved drugs from Japan, the USA and Europe; targets on 7,335 entries; 2026-edition ATC derivation | Mechanism around each edge: targets, metabolizing enzymes, efficacy and classification |
| WHO ATC/DDD Index | One classification row per substance or group level | Official 2026 classification with Defined Daily Doses across 14 anatomical groups, five levels deep | Classification parity across markets, with EMA catalogue ATC columns extending it to roughly 2,700 EU-authorised medicines |
| field | side | definition | example |
|---|---|---|---|
| ddi_pair | interaction | Queried drug entry paired with its counterpart compound or drug; one row per pair. | dr:D00564 x cpd:C00304 |
| class_flag | interaction | Grade carried on every pair: contraindication or precaution. | CI |
| class_label | interaction | Coarse grouping label on the pair; frequently unclassified, which is why the enrichment columns exist. | unclassified |
| TARGET | mechanism | Therapeutic target with human gene cross-reference; carried on 7,335 of 12,896 entries. | VKORC1 [HSA:79001] |
| METABOLISM | mechanism | Metabolizing enzymes and transporters acting on either side of the pair. | Enzyme: CYP2C9 [HSA:1559] |
| REMARK.ATC | classification | ATC classification code derived in the parent record, aligning the pair to the global vocabulary. | B01AA03 |
| drug_interactions | regulation | Verbatim interaction section quoted from the US structured product label of either product. | <label interaction wording> |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
KEGG REST API (DDI and Link Operations)
ddi_result_rows · link_pairs · conv_pairs
KEGG DRUG Database
RxNorm / RxNav Browser and APIs
rxcui · name · synonym …+6 more
openFDA Drug APIs
WHO ATC/DDD Index
ChEMBL Web Services API
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What counts as one record in a drug interaction database?
One pairwise assertion: a queried drug entry, a counterpart compound or drug, a contraindication or precaution flag and a coarse class label. A single drug with many partners appears once per pair, so roll up on the drug and scope down on the counterpart depending on the question.
How complete is drug interaction coverage?
The corpus spans all 12,896 approved-drug entries across Japan, the USA and Europe, though the pairs are drawn mainly from Japanese label material. For a US-facing product, pairing the edges with roughly 262,000 American structured product labels is what makes the combined output defensible.
Do the flags say how serious an interaction is?
They separate contraindication from precaution and stop there - no magnitude, and the class label is frequently unclassified. Teams needing graded severity layer it themselves from adverse-event frequency, boxed-warning presence or shared-target plausibility, each available as its own slice in the same pool.
Which dataset answers a drug interaction question?
The KEGG DDI corpus carries the edges themselves; the KEGG DRUG Database attaches mechanism - targets on 7,335 entries, metabolizing enzymes down to the CYP line; openFDA Drug APIs supply verbatim US label wording; RxNorm resolves names; WHO ATC/DDD standardizes classification. Delivered together, they arrive joined rather than merely adjacent.
Can interaction data be joined to a product catalog?
Yes. The corpus accepts US and Japanese marketed-product codes natively, and RxNorm's RXCUI bridge covers everything else, so a shelf item or claims code lands on the same identity the interaction pair describes - no proprietary crosswalk required.
How is drug interaction data delivered?
As typed rows - API, files, or straight into your warehouse - daily, weekly, or hourly, your call. Every delivery pins a release stamp, and retained consecutive deliveries turn the corpus into longitudinal history of pairs changing. Request a sample cut to your molecules and classes first.