PubChem PUG REST API - Chemical Structures & Properties

Datadory delivers pubchem pug rest api chemical structures properties data covering the largest public record of small-molecule chemistry: roughly 115 million compounds, 300 million substances and 900 million bioactivity data points, each compound carrying a 35-field dictionary - formula, SMILES and InChI identifiers, IUPAC name, XLogP, TPSA, hydrogen-bond and rotatable-bond counts, stereo counts, patent and literature links. Delivered as files, warehouse writes or API pulls, daily, weekly or hourly.

What is the PubChem PUG REST API - Chemical Structures & Properties dataset?

The reference shelf for small-molecule chemistry, filed as one addressed dataset inside Datadory's life sciences tools & services slice. Behind it sits NCBI's PubChem: roughly 115 million compounds, 300 million substances and 900 million bioactivity data points, organized on a clean three-key entity model - CID for compounds, SID for substances, AID for assays - that also reaches gene, protein, pathway, taxonomy and cell entities.

Thirty-five verified fields define the compound row: identity and nomenclature (formula, three identifier systems, generated IUPAC name), computed descriptors (XLogP, exact and monoisotopic mass, topological polar surface area, complexity, charge, hydrogen-bond and rotatable-bond counts), stereochemistry tallies, 3D conformer geometry, a substructure fingerprint, plus synonym, patent-count and literature-count layers. Eight response shapes ride along, JSON through CSV to SDF structure files and PNG-rendered depictions.

It scores 10 out of 10 on Datadory's rubric against a catalog averaging 7.81 across 1,744 datasets - a mark only 145 records earn, and it plays the compound half beside UniProt REST API's protein half. Get a sample of this dataset cut to your own compound list.

What do sample rows from the dataset look like?

One standardized compound per row, identity pinned first and computed values riding behind. Three everyday molecules, straight from the cataloging capture:

# delivered grain: one compound record per row

CID             : 2244
MolecularFormula : C9H8O4
MolecularWeight  : 180.16
IUPACName        : 2-acetyloxybenzoic acid

CID             : 2519
MolecularFormula : C8H10N4O2
MolecularWeight  : 194.19
IUPACName        : 1,3,7-trimethylpurine-2,6-dione

CID             : 1983
MolecularFormula : C8H9NO2
MolecularWeight  : 151.16
IUPACName        : N-(4-hydroxyphenyl)acetamide

Read the top row as three decisions already made upstream. Aspirin arrives as CID 2244 - the stable numeric address every other table joins on - with its formula already normalized to Hill notation and its weight carried to two decimals, not parsed out of prose. The caffeine row proves the dictionary holds heteroatom-rich ring systems as cleanly as substituted benzoic acids, and the paracetamol row shows the generated IUPAC name doing its job where marketing names diverge.

A fourth captured layer shows the annotation side. The same CID 2244 carries Title Aspirin alongside a description line attributed to its contributing source - the California Office of Environmental Health Hazard Assessment - so quoted text travels with provenance attached instead of anonymized. Synonym lists behave the same way: trade names, CAS numbers and systematic names arrive as an array, ready to match against whatever inventory list you hold.

Every value above comes from the catalog's August 2026 research pass, not a mock-up.

Which fields does the dataset include?

Thirty-five documented fields define the compound row, each definition verified during the August 2026 research pass - a bar only 145 of 1,744 cataloged datasets clear outright. Around fifty property tags exist in total; the remaining descriptor tags fold into your sample on request alongside derived columns such as Lipinski-style rule-of-five flags built from the counts below. They split into six jobs:

  • Identity: CID anchors every row, joined by Title, MolecularFormula and MolecularFormulaNoCharge.
  • Structure representations: SMILES with stereo and isotope layers, ConnectivitySMILES without them, standard InChI, the 27-character InChIKey hash and the base64-encoded Fingerprint2D substructure fingerprint.
  • Nomenclature: generated IUPACName plus the Synonym array - trade names, CAS numbers, systematic names in one field.
  • Mass and lipophilicity: MolecularWeight, ExactMass, MonoisotopicMass and computationally generated XLogP.
  • Drug-likeness descriptors: TPSA polar surface area, Complexity, formal Charge, HBondDonorCount, HBondAcceptorCount, RotatableBondCount and HeavyAtomCount.
  • Stereo, 3D shape and linkage: atom- and bond-level stereo counts, defined versus undefined splits, CovalentUnitCount, Volume3D, FeatureCount3D, ConformerCount3D, ConformerModelRMSD3D, then PatentCount, PatentFamilyCount, LiteratureCount and source-attributed Description text closing the record.

The full dictionary follows in tabular form below.

What does coverage look like across geography, time and granularity?

  • Geography - global by construction: chemical deposits and consolidated literature links arrive from sources worldwide, and no compound carries a jurisdiction flag. Region-shaped questions resolve by joining the patent and literature counts outward, not by filtering rows directly.
  • Temporal - compounds reach back through historical deposits, with assay deposition running continuously since PubChem's 2004 launch and the corpus refreshed continuously since. Records are current-state: today's identifiers, today's calculated values, today's link counts.
  • Granularity - four levels deep: whole compounds keyed by CID, deposited substances keyed by SID beneath them, individual assay results keyed by AID, and gene, protein, pathway, taxonomy and cell entities hanging off the same spine. Whole-corpus cuts (every compound under 500 Da with fewer than five hydrogen-bond donors) and narrow ones (one molecule's full synonym and assay neighborhood) both resolve without touching non-matching rows.

Set against the wider catalog - 1,744 datasets averaging 7.81 - the 10/10 prices in verified definitions, three clean join keys and a 35-field dictionary that stays flat at scale. Where it ranks: best life sciences tools services datasets.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

You pick the channel and cadence; normalizing the identifier graph, keeping standardized compounds apart from raw deposits and flattening fifty-plus descriptor tags into joinable columns stay our problem. Rows arrive CID-keyed, so this month's delivery appends cleanly onto last month's and joins whatever assay, target or inventory tables you already hold.

Hourly suits screening pipelines triaging newly deposited candidates while they are still decisions rather than history. Daily fits competitive monitoring of patent and literature link movement across a compound set. Weekly matches the rhythm most cheminformatics workflows actually run at. Whichever you choose, the thirty-five-field dictionary travels unchanged. A sample ships first either way - real rows for your own molecule list before any commitment.

Who uses this data, and for what?

Six jobs the compound shelf settles outright:

  1. Candidate triage before synthesis - filter 115 million compounds on XLogP, TPSA, hydrogen-bond donors and rotatable bonds to rank-order molecules worth ordering before anyone books instrument time; see ml model training for the model-building variant.
  2. Structure lookup and identifier resolution - turn names, CAS numbers or registry synonyms into canonical CIDs, SMILES and InChIKeys, or run the reverse; see api integration.
  3. Inventory reconciliation - match an internal catalog against canonical structures on InChIKey, surfacing duplicates, stale registrations and mislabelled stock without name-based guesswork.
  4. Virtual screening and similarity work - the substructure fingerprint plus stereo-aware identifier layers make neighbor searches a computed column rather than a chemistry engine you have to stand up.
  5. Market and landscape sizing - compound, substance and assay counts frame cheminformatics and compound-management market sizings; the same job serves the wider market sizing shelf.
  6. Citation-grade verification - quote formula, weight and generated names straight from the reference record when a paper or filing needs a number that holds up; see citation grade research.

Each job maps to a persona below, and the sample validates whichever one you came for.

Which personas get the most value?

Data Scientists & ML Engineers lead fit at relevance 3: decades of standardized structures with computed descriptors attached is the combination molecular-property models rarely get in one place - see life sciences tools services data for data scientists. Developers & Data-Product Builders also tag relevance 3, wiring stable CIDs into chemistry lookups and lab software where a number survives redesigns far better than a name ever will - see life sciences tools services data for developers builders. Journalists, Academics & Students (relevance 3) cite canonical structures and generated names with provenance attached. Competitive Intelligence & Product Teams and Market Researchers & Consultants (both relevance 1) read deposition and link counts as activity signals and market denominators.

Persona fit has edges worth naming: this describes structures; it does not measure expression, toxicity outcomes or clinical progress. Those live in neighboring catalogs - measured activities at ChEMBL, transcriptomic evidence at NCBI GEO, trial evidence at ClinicalTrials.gov. Pair accordingly rather than expecting one feed to cover the whole discovery chain.

What should I know before requesting a sample?

Three honest edges. First, standardized versus deposited: the 115-million compound count is deduplicated CIDs, while the roughly 300 million substance records count deposits as submitted - one molecule can carry many SIDs, so pick the grain that answers your question rather than the bigger number. Second, calculated means calculated: XLogP, TPSA and the descriptor family are model-generated values, not measurements, and any workflow needing measured biology should pair this record with a curated-activity source. Third, current-state only: trendlines over link counts or descriptor revisions come from repeated captures on a fixed schedule - which is what the daily, weekly or hourly cadence exists to automate. None of this is hidden at delivery; say which edge bites your question and the sample comes back shaped around it.

Which cards pair with this dataset?

Cards worth opening next:

Then get a sample of this dataset - real rows cut to your own molecule list before any commitment.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - the 35 fields of the compound record schema
fieldtypedefinitionexample
CIDintegerPubChem Compound ID - the unique identifier of the standardized compound record returned in every property-table row.2244
MolecularFormulastringMolecular formula in standard Hill notation element order, with total formal charge at the end.C9H8O4
MolecularFormulaNoChargestringMolecular formula, elements only, without total formal charge.
MolecularWeightnumberSum of all atomic weights of the constituent atoms in g/mol, assuming natural isotope abundance unless explicitly labelled.180.16
SMILESstringSMILES string including both stereochemical and isotopic information.
ConnectivitySMILESstringConnectivity-only SMILES string, without stereochemistry or isotope layers.
InChIstringStandard IUPAC International Chemical Identifier for the compound.
InChIKeystringHashed version of the full standard InChI, 27 characters.BSYNRYMUTXBXSQ-UHFFFAOYSA-N
IUPACNamestringChemical name systematically generated according to IUPAC nomenclature.2-acetyloxybenzoic acid
TitlestringThe title used for the compound summary page.Aspirin
XLogPnumberComputationally generated octanol-water partition coefficient, a measure of hydrophilicity/hydrophobicity.
ExactMassnumberMass of the most likely isotopic composition for a single molecule.
MonoisotopicMassnumberMass calculated using the most abundant isotope of each element.
TPSAnumberTopological polar surface area computed by the Ertl et al. algorithm.
ComplexitynumberMolecular complexity rating computed using the Bertz/Hendrickson/Ihlenfeldt formula.
ChargeintegerTotal (net) formal charge of the molecule.
HBondDonorCountintegerNumber of hydrogen-bond donors in the structure.
HBondAcceptorCountintegerNumber of hydrogen-bond acceptors in the structure.
RotatableBondCountintegerNumber of rotatable bonds.
HeavyAtomCountintegerNumber of non-hydrogen atoms.
AtomStereoCountintegerTotal number of atoms with tetrahedral (sp3) stereo.
DefinedAtomStereoCountintegerNumber of atoms with defined tetrahedral (sp3) stereo.
UndefinedAtomStereoCountintegerNumber of atoms with undefined tetrahedral (sp3) stereo.
BondStereoCountintegerTotal number of bonds with planar (sp2) stereo.
CovalentUnitCountintegerNumber of covalently bound units.
PatentCountintegerNumber of patent documents linked to the compound.
PatentFamilyCountintegerNumber of unique patent families linked to the compound.
LiteratureCountintegerNumber of articles linked to the compound by PubChem's consolidated literature analysis.
Volume3DnumberAnalytic volume of the first diverse conformer (default conformer) for the compound.
FeatureCount3DintegerTotal number of 3D features: sum of FeatureAcceptorCount3D, FeatureDonorCount3D, FeatureAnionCount3D, FeatureCationCount3D, FeatureRingCount3D and FeatureHydrophobeCount3D.
ConformerCount3DintegerNumber of conformers in the conformer model for the compound.
ConformerModelRMSD3DnumberConformer sampling RMSD in angstroms.
Fingerprint2DstringBase64-encoded PubChem Substructure Fingerprint of the molecule.
SynonymtextArray of synonyms for the compound (trade names, CAS numbers, systematic names).aspirin, ACETYLSALICYLIC ACID, 50-78-2
DescriptiontextTitle/description text with DescriptionSourceName and DescriptionURL identifying the contributing source.Aspirin can cause developmental toxicity and female reproductive toxicity according to an independent committee of scientific and health experts.

Questions buyers ask

How big is pubchem pug rest api chemical structures properties data?

Roughly 115 million standardized compounds keyed by PubChem CID, plus about 300 million deposited substance records keyed by SID and around 900 million bioactivity data points keyed by AID. The same spine reaches gene, protein, pathway, taxonomy and cell entities, so one compound identifier resolves outward across all of them.

What separates a substance record from a compound record?

Substances are deposits exactly as contributors submitted them, each with an SID. Compounds are the standardized small molecules extracted from those deposits, each with a CID. Every aspirin deposit converges on one aspirin CID, which is what makes compound-level analysis deduplicated and stable while substance-level analysis preserves provenance.

Which calculated properties ship on each compound?

Around fifty tags are documented; this dictionary details thirty-five. They cover formula and three weight measures, three structure representation systems, generated IUPAC naming, lipinski-adjacent descriptors such as XLogP, TPSA and hydrogen-bond and rotatable-bond counts, stereo tallies, 3D conformer geometry, a substructure fingerprint, and synonym, patent and literature link layers.

Does the dataset keep version history per compound?

No. Records mirror the present standardized state - current identifiers, current computed values, current link counts. History comes from scheduled captures appended snapshot by snapshot keyed on CID, which is precisely how requested deliveries are structured.

Can compound rows join to assay, target and pathway tables?

Yes, on identifiers rather than name matching. CIDs cross-reference SIDs, assay AIDs, gene IDs, protein accessions, pathway and taxonomy entities within the same corpus, so a compound column joins to measured bioactivity and target tables without fuzzy string reconciliation.

How fresh are the calculated descriptors?

Descriptor values recompute as the underlying structures and tooling evolve, and delivered rows reflect the state at capture time. Because deliveries key on CID and carry their capture date, a re-delivery refreshes values in place while appended snapshots preserve the earlier readings for anyone tracking drift.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing