Biotechnology Data Provider: 24 Cataloged Datasets · Head-to-head
ChEMBL vs ClinicalTrials.gov
Which biotechnology data provider: 24 cataloged datasets data fits your job: ChEMBL, or ClinicalTrials.gov. API, files, or your warehouse. Daily, weekly, or hourly.
ChEMBL
ClinicalTrials.gov
Coverage, side by side
| ChEMBL | ClinicalTrials.gov | |
|---|---|---|
| Geographic | Global literature-derived; source documents span international journals and patents, with no geographic columns on records | Global: studies registered from 200-plus countries; site locations down to facility, city and country |
| Temporal | Compounds from 1940s-era literature onward, curated through the current ChEMBL 37 release (2026) | Registry began 2000; records span 1999-present |
| Granularity | Per-molecule, per-target, per-assay, per-activity-record | One record per registered study with protocol, arms and interventions, outcomes, sites and results sections |
What each contains
They tie on 3 attributes. Pick by fit, not by loyalty.
| ChEMBL | ClinicalTrials.gov | |
|---|---|---|
| Steward | EMBL-EBI (European Bioinformatics Institute) | US National Library of Medicine, NIH |
| Subject lens | Manually curated bioactivity of drug-like molecules: compounds, targets, assays and activity measurements | Registry and results database of clinical studies: protocols, sponsors, phases, outcomes, sites |
| Geography | Global literature-derived; source documents span international journals and patents, with no geographic columns on records | Global: studies registered from 200-plus countries; site locations down to facility, city and country |
| Temporal reach | Compounds from 1940s-era literature onward, curated through the current ChEMBL 37 release (2026) | Registry began 2000; records span 1999-present |
| Granularity | Per-molecule, per-target, per-assay, per-activity-record | One record per registered study with protocol, arms and interventions, outcomes, sites and results sections |
| Documented fields | 27 | 16 |
| Definition confidence | Verified | Verified |
| Rubric rating | 10 out of 10 | 10 out of 10 |
| Delivery | Delivered daily, weekly, or hourly - your call | Delivered daily, weekly, or hourly - your call |
Or take both in one feed
They stack as candidate plus evidence. Mine ChEMBL for compounds clearing a potency bar against your target - pchembl_value above threshold, few RO5 violations - then interrogate the clinical record for what became of that chemotype: which studies exist, at which phases, under whose sponsorship, recruiting in which countries. Run it the other way too: a phase 2 study in the registry implies compounds worth looking up by name in the curated library, where max_phase and first_approval say whether the class ever crossed the finish line.
Join discipline, not luck: a molecule-level max_phase is the ceiling across a compound's whole history, never equivalent to any single study's declared phase; match on normalized names rather than trusting pref_name against unstructured intervention strings; and remember only one side knows geography. Datadory delivers both records reconciled to the same pipeline slice - delivered daily, weekly, or hourly, your call. Sample the pair and keep whichever answers faster. Or take both in one feed.
API, files, or your warehouse. Daily, weekly, or hourly.
Fair questions
Is ChEMBL better than ClinicalTrials.gov?
Different altitudes of the same pipeline. ChEMBL curates 24.5 million activity records binding molecules to targets in assays, with structures, calculated properties and standardized pchembl potency. ClinicalTrials.gov registers 599,549 studies with protocols, sponsors, phases, enrollment and results across 200-plus countries. Discovery questions favor ChEMBL; clinical questions favor the registry. Many teams keep both.
Which dataset covers more geography?
ClinicalTrials.gov, decisively. Studies register from more than 200 countries and each record lists sites down to facility, city and country, with coordinates where provided. ChEMBL is literature-derived - journals and patents from around the world, but no geographic columns on the records themselves.
Which goes back further in time?
ChEMBL. Its sources reach compounds from the 1940s onward, curated through the current ChEMBL 37 release. The registry began in 2000 and its records span 1999-present. For pre-2000 chemistry only ChEMBL remembers; once a molecule enters human testing, the registry takes over the story.
Do the two datasets describe drugs the same way?
No, and conflating them distorts models. ChEMBL describes a molecule: structure, calculated properties, assay potency, plus molecule-level flags such as max_phase and withdrawn_flag. ClinicalTrials.gov describes a study: one protocol, one sponsor, one phase, one set of sites. A compound's max_phase is the highest phase any of its studies reached - never equal to a single record's declared phase.
Whose field definitions are better documented?
Dead heat. Field definitions on both sides were captured as verified against live documentation - a bar much of the catalog misses - and both records score 10 out of 10 on Datadory's rubric. The difference is scope: 27 documented fields on the molecular side against 16 module-backed fields on the study side.
Which should a data scientist sample first?
Sample both, pick by fit. Modeling activity or selecting compounds starts with ChEMBL's activity records and property columns. Pipeline, competitive or site-level questions start with ClinicalTrials.gov's status, sponsor and location fields. Teams building bench-to-bedside views usually end up keeping both in rotation.