Datadory notebook
Pharmaceuticals Data Guide: Every Dataset, Source and License That Matters in 2026
Datadory delivers pharmaceuticals data guide with comprehensive historical coverage, validated schemas, and standardized fields — delivered daily, weekly, or on demand.
1,744 datasets. Pick your catch.
The pharmaceuticals data landscape in 2026
Of the 1,744 datasets Datadory catalogs across every industry, 22 sit in the pharmaceuticals pool — 13 primary pharmaceutical datasets and 9 adjacent records from neighboring industries such as biotechnology and drug retail. The defining feature of this pool is who publishes it: government regulators and public-science repositories dominate, and commercial vendors barely register.
Life-science content comes from EMBL-EBI's ChEMBL — 24,527,044 activities across 2,921,148 compounds and 18,552 targets — NCBI PubChem's compound and assay records, and KEGG's 12,896 approved-drug entries spanning Japan, the USA and Europe. NLM's RxNorm/RxNav suite normalises US drug names to RXCUIs, while the WHO Collaborating Centre's ATC/DDD Index sets the global classification standard across all 14 anatomical groups.
What are the core pharmaceuticals datasets?
Government and regulatory sources
Public-science and life-science sources
ChEMBL Web Services API (quality score 9) serves the full EMBL-EBI database — 24,527,044 activities, 2,921,148 distinct compounds, 3,824,604 compound records, 18,552 targets and 101,100 publications as of release 37 (2026-05-01) — under commercial delivery terms-SA 3.0 with no authentication, in JSON, XML, YAML, SDF or SVG, plus cheminformatics similarity and substructure search. PubChem PUG REST API (score 9, commercial delivery terms) adds property retrieval, synonym lookup, structure search and image rendering over tens of millions of compounds, backed by an FTP site publishing daily increments, weekly dumps on Sundays and monthly dumps on the first.
Standards and vocabularies
Commercial and proprietary sources
Pharmaceuticals data by use case: what do teams do with it?
Every major pharmaceuticals workload maps to named sources in this pool:
Data Coverage and Field Structure
Datadory delivers this dataset fully structured and normalized, ready for direct analysis without manual pipeline configuration or schema parsing.
Who uses pharmaceuticals data?
Four kinds of team account for most pharmaceuticals data work, and each arrives with its own search vocabulary.
Market researchers and competitive-intelligence analysts track approval pipelines rather than molecules: the EMA Medicines Catalogue's 2,730 authorised medicines and its companion tables reveal who gained centralised authorisation, which products hold orphan designations, and where shortages create openings. Searches look like "ema approved medicines list excel"; the market researchers page frames this pool through that lens.
How do you build a pharmaceuticals data stack?
A working pharmaceuticals stack layers five capabilities in order, because each step depends on the one before it.
What the numbers say
Six figures summarise the pool as of August 2026. Quotable version: Datadory catalogs 22 datasets for pharmaceuticals; ten of the 13 primaries are free outright and 6 expose official APIs; the FDA alone accounts for ~20.7 million adverse-event records.
| Figure | What it measures | Source |
|---|---|---|
| 22 | Datasets in the pharmaceuticals pool (13 primary + 9 secondary) | Datadory catalog |
| ~20.7M | Adverse-event report records (events from Q3 2004 onward) | openFDA Drug APIs |
| 2,730 | Centrally authorised medicines (2,337 human, 393 veterinary) | EMA Medicines Catalogue, Aug 2026 |
| 24,527,044 | Bioactivity records across 2,921,148 compounds and 18,552 targets | ChEMBL Web Services API |
| 12,896 | Approved-drug entries spanning Japan, USA and Europe | KEGG DRUG Database |
Keep reading
Continue through the pharmaceuticals cluster:
| Figure | What it measures | Source |
|---|---|---|
| 22 | Datasets in the pharmaceuticals pool (13 primary + 9 secondary) | Datadory catalog |
| ~20.7M | Adverse-event report records (events from Q3 2004 onward) | openFDA Drug APIs |
| 2,730 | Centrally authorised medicines (2,337 human, 393 veterinary) | EMA Medicines Catalogue, Aug 2026 |
| 24,527,044 | Bioactivity records across 2,921,148 compounds and 18,552 targets | ChEMBL Web Services API |
| 12,896 | Approved-drug entries spanning Japan, USA and Europe | KEGG DRUG Database |
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Is ChEMBL data free for commercial use?
Yes. EMBL-EBI publishes ChEMBL under commercial delivery terms-SA 3.0, which permits commercial use with attribution and share-alike, and the Web Services API requires no authentication. It serves 24,527,044 activity records across 2,921,148 compounds and 18,552 targets as JSON, XML, YAML or SDF, with quarterly database releases.