Datadory notebook
Cleaning product ingredients data: the formulation records behind what is inside the bottle
Datadory delivers household products data covering the whole cleaning-product ingredient stack: about 28,000 branded formulation records tying named products to ingredient tables with CAS numbers and percent concentrations, 3.8 million extracted chemical-by-product rows with weight fractions across 414,863 products and 40,459 unique substances, hazard dossiers for 1,376,722 chemicals and roughly 4,968 certified-product entries - delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
What is cleaning product ingredient data?
The search phrase asks where ingredient lists live; the working question underneath is who turns them into rows a team can query. On Datadory's household-products shelf four records split the job cleanly, and each ships as a product with sample rows, a field dictionary and a coverage statement:
- Consumer Product Information Database (CPID) owns the branded-formulation column: about 28,000 consumer products, each record tying a named product - UPC, physical form, stated usage, manufacturer address - to its own ingredient table.
- EPA Chemical and Products Database (CPDat) owns bulk composition: 3.8 million chemical-by-product rows carrying weight fractions across 414,863 products and 40,459 unique substances.
- EPA CompTox Chemicals Dashboard owns the dossier: identity, physicochemical properties, exposure estimates and hazard results for 1,376,722 substances, keyed on the same DTXSID spine.
- EPA Safer Choice Certified Products Database owns the certification bar: roughly 4,968 entries whose formulas cleared an ingredient-safety standard set by the program rather than self-declared on a label.
Density outruns headcount on this shelf. Household Products holds just seven primary assignments in Datadory's 1,744-dataset catalog and still lands three federal programs scoring 9 or better plus the slice's only quality-10 record, because no statistical agency runs a survey called household products: EPA certifies the safer formulations, BLS prices them at the register, EPA's toxicology programs map every substance that ever touched a label, and one private publisher reads the backs of boxes. Pooling four unrelated mandates onto one typed spine is the entire trick - which is why the useful question is not where the lists live, but which grain answers yours.
What do cleaning product ingredient rows look like?
Rows exactly as they land - two grains joined by one key. First the branded product grain, one record per dated formulation:
record_id : 19-006-143
product_name : Simple Green Heavy-Duty BBQ & Grill Cleaner, Aerosol,
Professional Use-02/23/2021
classification : Preparation
category : Inside the Home market : US/Canada
upc : 043318600142 form : aerosol
usage : BBQ and grill cleaner and degreaser
manufacturer : Sunshine Makers, Inc., Huntington Beach CA
ph_levels : 10.0-11.5 sds_date : February 23, 2021Underneath sits the ingredient grain - the composition table for that same record:
chemical : Water cas_no : 007732-18-5 pct_conc : >82
function : Diluent coc : No
chemical : Butane cas_no : 000106-97-8
function : Propellant coc : Yes
chemical : C9-11 Pareth-3 cas_no : 068439-46-3
function : Nonionic Surfactant coc : NoRead together, that is a retailer's shelf tag, a regulatory filing and the safety paperwork flattened into joinable rows. Three things repay attention. The coc column does real screening work: butane lights up as a chemical of concern while water and the surfactant pass - the whole hazard filter in one boolean. pct_conc arrives exactly as the label states it, thresholds included, so >82 is a claim to model as a band rather than a number to average. And the bulk grain adds what a plain ingredient list never carries:
# one row per chemical per product document
dtxsid : DTXSID7020182
curated_cas : 0000077-92-9
cleaned_weight_fraction : <decimal between 0 and 1>
weight_fraction_type : Reported # or Predicted (model estimated)
functional_use_category : Cleaning agent # OECD-harmonized class
puc : Cleaning products > All-purpose cleanerA presence list says a chemical is in the bottle; a weight-fraction row says how much, whether that figure was reported off the source document or predicted by a model, and which OECD-harmonized job the substance does in the formula. Decide your unit of analysis before the first join - brand, formulation or ingredient - because a naive row count mixes cleaners with molecules.
Which dataset covers branded cleaning product formulations?
Consumer Product Information Database (CPID) is the branded register, compiled by DeLima Associates of McLean, Virginia, initiated in 1994 with support from the National Institute of Environmental Health Sciences, and linking over 28,000 consumer brands to health effects. Each entry is a dated formulation record carrying commerce identity on one side and chemistry on the other, across ten categories - Auto Products, Inside the Home, Home Maintenance, Personal Care, Pesticides, Pet Care, Home Office, Hobby/Craft, Landscaping/Yard, Commercial/Institutional - plus curated collections for green, nano, children's, hand sanitizer and surface-sanitizer products.
Coverage spans the USA and Canada plus eight European country scopes (DE, EE, ES, FR, IT, NL, UK, FI), with formulation records dated roughly 2001 through 2026. History rides in the rows themselves: every record carries its own formulation date alongside date_entered and date_verified, so watching a competitor reformulate is a group-by rather than an archaeology project.
The twenty-two-field dictionary divides into three jobs. Identity and commerce fields - product_name, record_id, classification, market, upc, usage, form, manufacturer_information - key the register onto retail scan files without fuzzy matching. Chemistry fields - chemical, cas_no, ec_no, pct_conc, function, chemical_of_concern - describe what is in the bottle and why it is there, with registry numbers ready to join against any chemical master or watch list. Safety-narrative fields - ph_levels, sds_date, hazard_statements, acute_chronic_health_effects, first_aid - carry what the label and Safety Data Sheet actually claim, per brand, per formulation date. One normalization matters before any external join: CAS values arrive zero-padded eleven characters wide (007732-18-5, not 7732-18-5).
How does bulk ingredient chemistry scale past the brand shelf?
EPA Chemical and Products Database (CPDat) is the volume answer. Four counts describe the corpus: 485,048 underlying source documents, 414,863 products, 3.8 million extracted chemical records and 40,459 unique chemicals - each chemical-by-product row carrying a weight fraction, so approximate concentration is recorded rather than inferred. The current release stands at v4.0, dated April 2025, on a knowledgebase first assembled in October 2017.
Two harmonization layers make the rows queryable at scale. Products map to harmonized product-use categories - 300,308 of the 414,863 products carry a standard Product Use Category link - so aggregation runs on how products are used instead of how brands spell their names. Chemicals carry functional-use and general-use records aligned to OECD-harmonized functions, which is why Cleaning agent sorts into one bucket whether the label says surfactant, wetting agent or detergent.
Provenance survives audit because every assignment discloses how it was made. Classification methods run from Manual through Batch Assign to Automated NLP-based tagging; product attributes tag physical form, exposed population and micro-environment; and every row chains back to a dated source document - manufacturer safety sheets, retailer data, government reports - with HERO references and, where pesticides or disinfectants are involved, EPA registration numbers. One honest boundary shapes expectations: the underlying documents skew toward United States commerce, so category prevalence reads American even where the chemistry travels globally.
Where does the ingredient row stop and the chemical dossier begin?
A composition row ends at the substance; the dossier starts there. EPA CompTox Chemicals Dashboard carries identity, physicochemical properties, environmental fate, exposure estimates and toxicity and bioassay results for 1,376,722 substances, each hanging off a stable DTXSID rather than a spelling. It reads both directions: look up a chemical and see which consumer product-use categories it appears in, or start from a category and see which chemicals surface. More than 300 curated chemical lists sit on top of the substance records, so whether a substance lands on a relevant list is a field, not a literature review; the current release badge reads v2.8.0, stamped October 2025.
EPA Safer Choice Certified Products Database supplies the bar those compositions get measured against: about 4,968 entries - 4,858 Safer Choice certifications plus 110 legacy Design for Environment records - covering roughly 1,506 distinct product names from 331 partner companies across around 50 product types. All-purpose cleaners, degreasers, dish soaps, laundry detergents, toilet bowl and window cleaners anchor the household shelf, and the register reaches commercial lines such as parts washers and medical instrument cleaners. Every row names the company, sector, UPCs and GTINs, partner-since year, fragrance-free and outdoor-use flags, and whether the partner is current on yearly review. One structural quirk matters before counting: an entry repeats once per sector/product-type combination, so deduplicate on product name for 1,506 distinct formulations or keep the listings to preserve the merchandising grain - both are legitimate tables answering different questions.
Run together they make the green-claim audit mechanical: the certification register says whose formulas cleared the standard, the dossier says what the ingredients are known to do, and the composition rows say how much of each actually sits in the category.
Who builds on cleaning product ingredient data?
- Regulatory and EHS teams screen formulations before they enter a facility or supply contract: filter on
chemical_of_concern, read the GHShazard_statements, checkph_levelsagainst handling limits, and the result is a defensible screening table instead of a folder of photocopied spec sheets. - Competitive-intel and formulation chemists benchmark rivals' recipes - which surfactant system a category leader uses, what propellant blend replaced butane and when, where concentration bands cluster - with dated records turning reformulation into a timeline rather than a rumor.
- E-commerce and catalog teams enrich SKU data with UPC-keyed ingredient and hazard attributes, so a product page states what is inside using the numbers the manufacturer filed.
- Market researchers size and price the category: BLS's price and inflation tools put monthly CPI items for household cleaning and paper products beside per-unit Average Price series - items beginning December 1997, averages reaching back to about 1980, across regions, nine divisions and 23 metro areas, only three of which publish monthly while the rest report every other month.
- Model builders train ingredient classifiers and recommendation models on DTXSID-keyed corpora with numeric targets, while vision teams take the cleaning-products occlusion image set as a scoped pilot: the card documents 15,000 annotated images across Lysol, Clorox, Mr. Clean and Vanish, and the committed public slice hosts five - a gap stated up front rather than discovered mid-project.
Why get cleaning product ingredients data through Datadory?
Files, feeds, or straight into your warehouse. Daily, weekly, or hourly - your call. Datadory handles the reconciliation upstream of you: normalized CAS encodings, units and thresholds kept beside values so >82 never masquerades as 82, product and ingredient grains delivered as separate typed tables with the join key intact, and every shipment carrying the same field dictionary, sample rows and coverage statement mapped to the records you named. The schema in the sample is the schema you ship against.
Start with a sample: name the brands, chemicals, categories and markets you need, and the extract arrives cut to them - real rows, in the exact shapes shown above.
Where to go next
The household-products data hub holds the pooled view of all seven primary assignments and shows where each record sits on the shelf. Head-to-head context lives in EPA CPDat vs CompTox Chemicals Dashboard for the composition-versus-dossier split and EPA Safer Choice vs BLS Price & Inflation Data Tools for the certification-versus-price pairing. Where the records rank is on 10 Best Household Products Datasets in 2026, scored on Datadory's rubric - request samples in the same schema and they land in one warehouse.
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
Consumer Product Information Database (CPID)
product_name · record_id · classification …+19 more
EPA Chemical and Products Database (CPDat)
DTXSID · DTXRID · Component …+3 more
EPA CompTox Chemicals Dashboard
DTXSID · CASRN
EPA Safer Choice Certified Products Database
program · category · sector …+11 more
BLS Price & Inflation Data Tools (CPI, Average Price)
series_id · series_title · item_code …+12 more
Household Cleaning Products Occlusion Image Dataset
file_name · quality · product_category …+2 more
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
How many cleaning product ingredient records are there?
Three counts carry the category. The Consumer Product Information Database holds about 28,000 branded formulation records, each with its own ingredient table. EPA's CPDat holds 3.8 million chemical-by-product rows with weight fractions, extracted from 485,048 source documents across 414,863 products and 40,459 unique substances. The Safer Choice register adds roughly 4,968 certified-product entries.
What is the difference between CPID and CPDat formulation records?
Grain and purpose. CPID publishes one record per branded product formulation - UPC, physical form, usage statement, manufacturer address, pH band and the safety narrative beside an ingredient table with CAS numbers and percent concentrations. CPDat publishes extracted chemical-by-document rows: weight fractions, harmonized product-use categories and OECD-aligned functional-use classes, built for aggregation rather than brand lookup. Brand-level searchability sits in the first; volume and harmonization sit in the second.
Can ingredient data flag chemicals of concern automatically?
Yes, twice over. Every CPID ingredient line carries a chemical_of_concern boolean - in the worked example, butane flags while water and the nonionic surfactant pass, which is a hazard filter in one column. At the substance level, the CompTox Chemicals Dashboard attaches memberships in more than 300 curated chemical lists to each DTXSID, so list status is a field rather than a literature review.
Can you track cleaning product reformulations over time?
Yes. CPID records are dated - formulation dates ride on the records from roughly 2001 through 2026, alongside date_entered and date_verified fields - so a reformulation is a group-by, not an archaeology project. CPDat carries a document date on every extraction, so cohort comparisons run on pinned snapshots of the paperwork rather than on assumptions about releases.
Which cleaning products hold Safer Choice certification?
About 4,968 entries - 4,858 Safer Choice certifications plus 110 legacy Design for Environment records - covering roughly 1,506 distinct product names from 331 partner companies across around 50 product types. All-purpose cleaners, degreasers, dish soaps, laundry detergents and toilet bowl cleaners anchor the household shelf, with commercial lines such as parts washers alongside. Entries repeat once per sector/product-type combination, so deduplicate on product name for distinct formulations.
How is cleaning product ingredient data delivered?
As typed rows - files, scheduled feeds, or straight into your warehouse, daily, weekly or hourly, your call. Product and ingredient grains arrive as separate tables with the join key intact, CAS encodings normalized, and every delivery carrying the same field dictionary, sample rows and coverage statement mapped to the records you named.