EPA CompTox Chemicals Dashboard

Datadory delivers epa comptox chemicals dashboard data: 1,376,722 chemical substances from the EPA Computational Toxicology Program's dashboard, each keyed by DTXSID and CASRN with preferred name, molecular formula, monoisotopic mass, physicochemical properties, CPDat-derived product-use categories, exposure estimates, hazard and bioassay results and 300+ curated chemical-list memberships, delivered daily, weekly, or hourly.

What is the EPA CompTox Chemicals Dashboard?

It is the U.S. EPA Computational Toxicology Program's answer to a question every household-products team eventually asks: what is actually known about this chemical - not one hazard number, but the whole dossier. The dashboard's landing page advertises 'Search 1,376,722 Chemicals' across three navigation tabs - Chemicals, Products/Use Categories and Assay/Gene - and per EPA's own tool description it integrates physicochemical properties, environmental fate and transport, exposure, usage, in vivo toxicity and in vitro bioassay data for over one million chemicals.

Three structural facts separate it from the ingredient lists most product teams work from. First, it is substance-first: one record per chemical substance, keyed by the DSSTox Substance Identifier (DTXSID), so every other fact hangs off a stable ID rather than a spelling of a name. Second, it is two-way on products: you can look up a chemical and see which consumer product categories it appears in, or start from a product category and see which chemicals show up - the bridge between what is in the bottle and what is known about the ingredient. Third, it is curated at the margins: more than 300 chemical lists, built on structure or category, sit on top of the substance records, so 'is this on a relevant list' is a field, not a literature review.

The current release badge reads v2.8.0, footer-stamped October 2025, and the record scores 9/10 on Datadory's rubric. Within the household products shelf it plays the hazard and exposure column: EPA CPDat tells you how much of a chemical sits in which products; this dashboard tells you what that chemical does once it gets out.

What do sample rows look like?

Two views carry the whole grammar. The first is a substance row as it lands in your warehouse - one row per chemical substance:

dtxsid              : DTXSID7020182
casrn               : 7732-18-5
preferred_name      : Water
molecular_formula   : H2O
monoisotopic_mass   : 18.0106
qc_level            : Verified by EPA

The second is the product-use edge - the reason this record lives on the household-products shelf rather than in a chemistry library. One row per chemical-per-product-category, CPDat-derived:

# product/use edge: one row per chemical x product-use category

dtxsid                  : DTXSID7020182
product_use_category    : Household cleaning products
exposure_tab_summary    : predicted + aggregate exposure estimates attached
assay_gene_hits         : in vivo tox + in vitro bioassay panels linked
list_memberships        : 300+ curated lists, flagged where present

Read them together and the join logic writes itself: dtxsid keys the substance, the product-use category keys the shelf, and everything else - properties, exposure, hazard, assay hits, list membership - arrives joined to the substance rather than looked up per project. The example above is water, the dashboard's canonical demo substance; the delivered table runs the same shape across all 1,376,722 records, with the full field set below replacing the six columns the web UI surfaces per substance.

Which fields does the dataset include?

Ten fields make up the dictionary, graded inferred on the source card - the dashboard renders per-chemical detail client-side, so definitions were mapped during the August 2026 research pass against the substance page and batch-search inputs rather than scraped from a schema document. They divide into four jobs:

  • Identity: DTXSID is the primary key and the join point to every other chemistry dataset in this catalog; CASRN is the identifier your suppliers' documents, safety data sheets and regulatory filings already use; Preferred Name gives one canonical display name per substance.
  • Structure: Molecular Formula supports formula-level screening, with MS-ready variants built for batch search; Monoisotopic Mass turns an LC-MS peak table into named substances without any fuzzy matching.
  • Behavior: the Physicochemical Properties panel - melting point, vapor pressure, logKow and their siblings - drives formulation, fate and bioavailability reasoning; Exposure Estimates summarize predicted and aggregate exposure per substance.
  • Context: Product Use Category ties the substance back to the shelf; Hazard / Bioassay Data carries in vivo toxicity plus high-throughput screening results from the Assay/Gene tab; Chemical List Membership flags which of the 300+ curated lists a substance appears on.

Field-level confidence matters when you build on it: treat the identity block as bedrock, and ask for the sample to confirm exactly which property endpoints your categories need before committing a pipeline.

What does coverage look like across geography, time and granularity?

Geography - global chemical scope with U.S.-focused regulatory and exposure context. The substance universe is worldwide chemistry; the exposure and usage layers lean American, because the underlying product-prevalence and exposure modeling is EPA work. A European formulation question gets full substance and hazard coverage but should expect the product-use prevalence to read US-shaped.

Temporal - continuously maintained, not versioned annually: the resource carries a running release badge (v2.8.0, October 2025 at research time) rather than a fixed publication calendar, and substance records are revised as the underlying DSSTox curation absorbs new literature. There is no deep time-series here - this is a current-state dossier per chemical, so trend work means snapshotting deliveries over time.

Granularity - one record per chemical substance, with related tables hanging off it: product-use categories, assays, genes and list memberships. That grain makes joins cheap and counts honest: filter to a category, count distinct DTXSIDs, and the result means substances, not listings - a sharper unit than most product-side registers, where one brand appears once per shelf it sits on.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

You pick the channel and the cadence; identifier resolution, MS-ready formula handling and list-membership normalization stay our problem. Substances arrive normalized to the dictionary above with product-use categories, exposure summaries and assay links joined on, so a delta pull appends newly curated substances instead of re-delivering all 1,376,722.

Hourly or daily suits monitoring builds that flag when a substance in your formulation portfolio picks up a new list membership or hazard signal. Weekly fits regulatory and compliance reviews reconciling a supplier's ingredient deck against current curation. Monthly-to-quarterly suits R&D screening, where the batch of candidate chemistries is stable and the dossier depth matters more than the clock. Whichever you choose, the dictionary travels unchanged and the sample ships first - real rows for the substances you actually name.

Who uses this data, and for what?

Five jobs this record settles outright:

  1. Ingredient screening and reformulation - resolve every candidate chemistry to a DTXSID, then let hazard, exposure and list-membership fields rank the shortlist before anyone books a toxicology consult.
  2. Regulatory and compliance readiness - check which of the 300+ curated lists a substance sits on, and whether its dossier has moved since the last review cycle, as evidence rather than anecdote.
  3. Mass-spectrometry and lab triage - push an MS peak list through formula and monoisotopic-mass batch search and get named substances back; see data enrichment use cases.
  4. Exposure and risk modeling - combine the physicochemical panel with predicted exposure estimates to prioritize which product categories deserve measurement effort.
  5. Competitive chemistry intelligence - map which substances concentrate in which product-use categories, then watch the edges move as curation absorbs new products; see competitive intelligence use cases.

The common thread: every job starts from a chemical someone already has - a supplier name, an INCI string, a lab peak - and this is the register that resolves it into a full dossier.

Which personas get the most value?

Data scientists and ML engineers get a substance-keyed corpus with stable identifiers and structured property panels - the kind of base table hazard prediction and claim-verification models actually want; see data scientists household products use cases. Regulatory and EHS leads get list membership and hazard context per ingredient, turning 'we think it is fine' into a queryable dossier; see procurement and EHS persona page. Product developers and formulators screen candidate chemistries against physicochemical behavior and product-use prevalence before the bench work starts. Market researchers and consultants read the product-use category edges as a map of where chemistry concentrates by shelf; see market researchers household products use cases. Journalists and academics cite a federal computational-toxicology program behind every substance claim; see journalists academics household products use cases.

What should you know before requesting a sample?

Three notes worth having upfront.

First, the dictionary is graded inferred, not verified. The dashboard renders per-substance detail client-side, so field definitions were mapped against the substance pages and batch-search inputs during the August 2026 research pass rather than read off a published schema. Identity fields are safe to build on immediately; confirm the exact property endpoints your use case needs when you request the sample.

Second, this is a current-state register, not a time series. Substance records are revised continuously under running releases - v2.8.0 in October 2025 at research time - so longitudinal claims need snapshots captured over time. Decide your snapshot cadence before the first delivery, not after the second.

Third, product-use prevalence is EPA-derived and therefore US-shaped. The substance and hazard layers travel globally; the category-prevalence layer reflects the American product documentation underneath it. Pair it with EPA CPDat for weight-fraction depth inside the bottle - they join cleanly on DTXSID - or head to Get a sample of this dataset and name your substances; real rows come back in exactly the schema shown above.

Where does it sit on the household-products shelf?

On the household-products data hub shelf this record owns the hazard-and-exposure column: 1,376,722 substances of context behind every formulation row the other records carry. Around it: EPA CPDat supplies the weight fractions inside the bottle, the Consumer Product Information Database adds branded ingredient tables with SDS pointers, EPA Safer Choice Certified Products Database marks which finished products cleared a federal standard, and BLS Price & Inflation Data Tools track what the category costs over time. Head-to-head on chemistry breadth versus product depth: CPDat vs CompTox Chemicals Dashboard. Seven records total, one shelf - request samples from any card.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - ten inferred-grain fields, one row per chemical substance with related product-use, assay and list tables
fieldtypedefinitionexample
DTXSIDstringDSSTox Substance Identifier - the dashboard's primary chemical key and the join point to every other chemistry dataset in this catalog.DTXSID7020182
CASRNstringCAS Registry Number, the identifier used across supplier documents, safety data sheets and regulatory filings.7732-18-5
Preferred NamestringDSSTox-preferred display name of the chemical - one canonical label per substance.
Molecular FormulastringChemical formula, with MS-ready variants maintained for batch search against instrument outputs.
Monoisotopic MassnumberExact monoisotopic mass usable as a batch-search input for resolving LC-MS peaks to substances.
Physicochemical PropertiestextPanel of measured and predicted properties shown per chemical - melting point, vapor pressure, logKow and related endpoints.
Product Use CategorystringCPDat-derived consumer and industrial product-use categories in which the chemical appears.Household cleaning products
Exposure EstimatestextExposure-tab summary of predicted and aggregate exposure data for the substance.
Hazard / Bioassay DatatextIn vivo toxicity and in vitro high-throughput screening bioassay results associated with the substance via the Assay/Gene tab.
Chemical List MembershiptextMembership among the 300+ curated chemical lists built on structure or category.

Questions buyers ask

How many chemicals does the CompTox Chemicals Dashboard cover?

The landing page advertises a searchable universe of 1,376,722 chemicals at the August 2026 research pass, arranged under three tabs - Chemicals, Products/Use Categories and Assay/Gene - with more than 300 curated chemical lists layered over the substance records.

Which identifiers key the chemical records?

Each substance carries a DSSTox Substance Identifier (DTXSID), the dashboard's primary key, plus its CAS Registry Number (CASRN). Batch search also accepts names, exact and MS-ready formulas, and monoisotopic mass, so mass-spectrometry hit lists resolve to named substances without manual lookup.

Does the dashboard connect chemicals to consumer products?

Yes. A Products/Use Categories view reports which consumer product categories a chemical appears in, derived from EPA's CPDat product-composition work, and each substance's Exposure tab carries the same product-use prevalence beside predicted and aggregate exposure estimates.

Can I run batch searches across many chemicals at once?

Yes - advanced and batch search accept lists of chemical identifiers, MS-ready formulas, exact formulas and monoisotopic masses, which is what makes the record useful for screening a candidate list or resolving an instrument output. Delivered tables preserve those inputs alongside the resolved substance keys so provenance never breaks.

Is there a time series of how chemical assessments changed?

No - the dashboard is a continuously maintained current-state register, release-badged v2.8.0 in October 2025 at research time, not an archive. Trend questions need snapshots: Datadory can deliver dated snapshots on a cadence you choose, so change detection between two dates becomes a simple diff on DTXSID.

How fresh is the data, and how do I keep it current?

The resource is continuously maintained rather than annually versioned, and Datadory mirrors that with deliveries daily, weekly, or hourly through API, files, or your warehouse. Delta pulls append newly curated substances and revised dossiers, so a monitoring build sees a new hazard or list signal the day after it lands upstream.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing