Datadory notebook

Is Chembl Data Free For Machine Learning Data: Dataset Structure and Field Coverage

Datadory delivers is chembl data free for machine learning data covering comprehensive field definitions, entity mappings, and historical time series — structured for direct analytics and delivered on demand.

1,744 datasets. Pick your catch.

What can you actually train on inside 24.5 million activity records?

Every activity row ties a molecule to a target through a specific assay drawn from a source document, which is the shape most property- and activity-prediction work needs. Because the database is literature-derived, coverage runs from compounds documented in the 1940s through current publications, sourced globally across international journals and patents.

That structure supports the hit-triage workflow Datadory tags for the industry: pull bioactivity records for one target from the 24,527,044 activities, then cross-reference candidate structures against PubChem's ~179.5 million compounds to widen the search space. It equally supports target-side models, because the 18,552 targets arrive with assay provenance rather than bare identifiers - see the bioactivity data terminology for how these fields fit together.

One caveat belongs in any model card. Curation is manual, so the label surface reflects the measurements scientists chose to run and publish, not a uniform screen. That sampling bias is worth stating explicitly when you report performance.

What should you check before shipping a model trained on ChEMBL?

Expect the counters to move, too. Every figure on this page - 2,921,148 molecules, 18,552 targets, 24,527,044 activities, 1,970,438 assays - came from the live API in August 2026 and will rise at the next quarterly drop, so quote them as of August 2026 rather than as permanent constants.

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Biotechnology

ChEMBL

Biotechnology

PubChem

Biotechnology

UniProt

Biotechnology

RCSB Protein Data Bank

Biotechnology Global drug content

DrugBank Online

Biotechnology Global reference content

KEGG

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Is ChEMBL data free for machine learning?

Yes. ChEMBL publishes 2,921,148 molecules, 18,552 targets and 24,527,044 activity records under commercial delivery terms-SA 3.0, with a REST API plus SQLite, MySQL, PostgreSQL and SDF dumps. The full dump runs roughly 204 MB compressed - small enough to load locally - and new releases arrive quarterly, with ChEMBL 37 issued in May 2026.

Is ChEMBL or PubChem better for machine learning?

They solve different problems. ChEMBL contributes manually curated labels - 24.5 million activities across 18,552 targets - under commercial delivery terms-SA 3.0. PubChem contributes breadth: CIDs up to ~179.5 million as commercial delivery terms with no attribution duty, though submitters may claim rights in portions of deposited data. Many pipelines join both.