Personal Care Products Data Provider: 11 Cataloged Datasets · Head-to-head

Open Beauty Facts — Global Cosmetics Product Database (Exports + Live API) vs Hugging Face Datasets — Cosmetics & Beauty Collections

Which personal care products data provider: 11 cataloged datasets data fits your job: Open Beauty Facts — Global Cosmetics Product Database, or Hugging Face Datasets — Cosmetics & Beauty Collections. API, files, or your warehouse. Daily, weekly, or hourly.

Personal Care Products Data Provider: 11 Cataloged Datasets Global volunteer contributions · Per-record created and last-modified timestamps reaching back to the project's start

Open Beauty Facts — Global Cosmetics Product Database (Exports + Live API)

Personal Care Products Data Provider: 11 Cataloged Datasets Global in the beauty split with France/Europe weighting observed in country tags · Product records from 2014 onward

Hugging Face Datasets — Cosmetics & Beauty Collections

Where the fields line up

7 shared fields — join on these.

Field Open Beauty Facts — Global Cosmetics Product Database Hugging Face Datasets — Cosmetics & Beauty Collections
code Barcode of the product - EAN-13 in most markets; products with no printed barcode get a number from the reserved 200 prefix, so every row still has a stable key. documented
product_name Product name exactly as entered by the contributor who photographed or scanned it - shelf naming, un-normalized, which is why it reads like a label rather than a catalog. documented
brands Brand name(s) on the packaging; a normalized lowercase brands_tags companion field holds the joinable versions for grouping and deduplication. documented
categories_tags Normalized hierarchical category tags, comma separated, walking from the broad shelf down to the specific product type. documented
ingredients_text Full ingredient list as printed on the packaging, INCI order preserved - the raw material for formulation analysis and ingredient-trend work. documented
states_tags Completion state of the record - markers like ingredients-to-be-completed or photos-uploaded that tell you how much of the row you can trust at a glance. documented
completeness Completeness score of the record between 0 and 1 - the project's own measure of how much of the schema a contributor filled in. documented

Coverage, side by side

Open Beauty Facts — Global Cosmetics Product Database Hugging Face Datasets — Cosmetics & Beauty Collections
Geographic Global volunteer contributions, browsable by country facet; no published country breakdown Global in the beauty split with France/Europe weighting observed in country tags; reviews drawn from a US marketplace
Temporal Per-record created and last-modified timestamps reaching back to the project's start, plus rolling 14-day delta windows Product records from 2014 onward; anchor repository last modified 2026-08-20; review rows timestamped individually

What each contains

Pick by fit, not by loyalty.

Open Beauty Facts — Global Cosmetics Product Database Hugging Face Datasets — Cosmetics & Beauty Collections
Publisher Open Beauty Facts - the cosmetics sibling of Open Food Facts, run by a French non-profit Hugging Face Hub - community-uploaded repositories, anchored by the official Open Beauty Facts export
Subject lens The product record itself: barcode, name, brands, categories, INCI ingredient text, labels, packaging, origins, manufacturing places, stores, countries, images Distribution plus adjacency: the same records re-cut as a columnar beauty split, flanked by Amazon beauty reviews and an ingredient-description table
Geographic coverage Global volunteer contributions, browsable by country facet; no published country breakdown Global in the beauty split with France/Europe weighting observed in country tags; reviews drawn from a US marketplace
Temporal coverage Per-record created and last-modified timestamps reaching back to the project's start, plus rolling 14-day delta windows Product records from 2014 onward; anchor repository last modified 2026-08-20; review rows timestamped individually
Detail level One row per product barcode (SKU level) with multi-value tag fields for categories, brands, ingredients and countries One row per barcode in the beauty split, one row per review, one row per ingredient across its three sets
Scale 73,522 beauty products (site counter, 2026-08-21) Beauty split pinned at 73,421 rows x 111 columns; roughly 700,000 Amazon reviews; a small ingredient table
Documented fields 31 12
Best for Formulation and compliance questions: ingredient screening, allergen and additive checks, label claims, packaging and manufacturing-place mapping Pipeline and feature questions: columnar reads over product barcodes, review sentiment beside products, ingredient knowledge lookups

What each does better

Open Beauty Facts

A Garnier hair mask keyed 3600542613316 arrives carrying its French label name and its complete ingredient chain verbatim.

Quality machinery on every row. Each record ships with a 0-1 completeness score, states_tags completion flags, data_quality_errors_tags warnings, the contributing creator and a unique_scans_n counter that turns app scans into a popularity signal. The OLAPLEX conditioner row demonstrates the ceiling: a fully populated INCI string from water through bis-aminopropyl diglycol dimaleate.

A census that keeps its history. Records carry created and last-modified timestamps back to the project's start, rolling 14-day delta windows sit on top of the cumulative snapshot, and a secondary household-products lens extends the shelf beyond beauty proper. Browse how such inventories behave in crowdsourced product database.

The only perfect rubric score on its shelf. Ten out of ten makes it the sole quality-10 record among the seven cataloged Personal Care Products datasets.

Hugging Face Datasets

Columnar convenience at scale. The beauty split lands as 73,421 rows by 111 columns in roughly 59 MB, light enough to load whole and narrow enough to read column by column when only a handful of fields matter. Same barcodes, friendlier grain for notebooks and training loops.

Two companions the original does not have. An Amazon beauty-review corpus of roughly 700,000 rows - star rating, title, text, ASIN, user, timestamp, helpful votes and verified_purchase flags - puts consumer sentiment next to product identity, and an ingredient-to-description table supplies plain-language write-ups (glycerin: an oldie but a goodie, in use for more than 50 years) that no product record contains.

One hub, adjacent shelves. A cosmetics search reaches past the non-profit census into whatever the community has uploaded - retailer catalogs, review corpora, ingredient encyclopedias - which makes the Hub side the wider net for competitive tracking work sketched in competitor tracking.

Where they're equivalent

More unites them than the two-point rubric gap suggests. Both document the same volunteer-contributed census at the same grain - one row per product barcode, multi-value category, brand and ingredient tag lists included - and the Hub's anchor repository is the official Open Beauty Facts export, so a given barcode resolves to the same product either way. Both field dictionaries are verified during research, and both carry the catalog's daily cadence tag.

They also share the census's design trade-offs. Coverage is global in principle but skews France and Europe in the observed country tags, and neither side publishes a country breakdown. Completeness is uneven by construction - many records sit flagged ingredients-to-be-completed - which is precisely why the 0-1 completeness score travels as a column everywhere the rows go. Counting products gives near-identical answers on both sides: 73,522 by the site counter on 2026-08-21, 73,421 rows pinned on the export side, the same shelf measured days apart.

The verdict

Verdict: sample both, pick by fit - one is the census, the other its distribution plus neighbors, so the question decides.

Reach for Hugging Face Datasets — Cosmetics & Beauty Collections when the question outgrows the product row: pipeline-friendly columnar reads over the same barcodes, review-level sentiment features joined to products, ingredient-knowledge lookups beside the catalog.

Data scientists usually reach for the second when features matter more than fidelity, and drop to the first when a claim needs sourcing; market researchers tend to run the order in reverse - the census first, the review corpus last. Competitive-intel product teams mostly want the product grid either way.

Sample both, pick by fit. See Open Beauty Facts — Global Cosmetics Product Database · See Hugging Face Datasets — Cosmetics & Beauty Collections

Or take both in one feed

Worked example: take the beauty split's 73,421 barcodes, enrich each with the columns the wide dictionary adds - labels, allergens, traces, additives, packaging, manufacturing places - then fan out to the review rows through an ASIN-to-barcode crosswalk so sentiment sits beside formulation. A Kiehl's moisturizer keyed 0000527000057 arrives on the export side with a 0.27 completeness score and to-be-completed flags; its shelf-mates on the review side arrive rated five stars with verified-purchase flags and helpful-vote counts attached.

Three seams decide whether the merge holds. First, keys: code matches exactly on the product grain, but product names differ in shape - flat strings on the original, {lang, text} lists on the export. Second, grain: products fan into reviews one-to-many, and the ingredient table joins on ingredient token rather than barcode, so aggregate before you weight. Third, versions: the export trails edits to the live record, so pin a snapshot date before diffing. The barcode key itself is profiled in barcode product key.

Browse the rest of the shelf at the personal care products data hub, or see how the pairing feeds ML model training. Datadory ships either record alone or merged onto one barcode key, delivered daily, weekly, or hourly - your call. Or take both in one feed.

API, files, or your warehouse. Daily, weekly, or hourly.

Fair questions

Is Open Beauty Facts better than Hugging Face Datasets?

Better at different jobs. Open Beauty Facts owns the product record: 31 documented columns spanning INCI ingredient text, labels, allergens and traces, additives, packaging, origins, manufacturing places and brand owner across 73,522 products, scoring 10 out of 10. The Hugging Face collections own convenience and adjacency: the same records as a 73,421-row, 111-column Parquet beauty split plus roughly 700,000 Amazon beauty reviews and an ingredient-description table, scoring 8 out of 10.

Are these two datasets actually different data?

Mostly no - and that is the interesting part. The Hub's anchor repository is the official Open Beauty Facts export, so the beauty split's barcodes resolve to the same volunteer-contributed products. What differs is shape and surroundings: 12 documented keys against 31, Unix timestamps against ISO-8601, and two companion tables - consumer reviews and ingredient descriptions - that exist only on the Hub side.

Which dataset has more fields, Open Beauty Facts or Hugging Face Datasets?

Thirty-one documented columns against twelve documented keys, and the gap is regulatory depth. The original spells out generic names, stated quantities with normalized gram equivalents, origins, manufacturing places, stores and countries, emb approval codes, labels and certifications, allergens, traces, additive counts, packaging, brand owner, data-quality errors, scan counts and images. The export keeps the core identity-and-formulation spine - barcode, name, brands, categories, ingredients, completeness, states, timestamps - and spends its spare keys on reviews.

Can the two datasets be joined?

Yes, on the barcode. The code field is shared and stable at the product grain, so enrichment is a straight merge: take the 111-column export split, add the wide dictionary's labels, allergens, additives and manufacturing places, then fan out to the roughly 700,000 review rows through an ASIN-to-barcode crosswalk. Mind the shapes - names arrive as strings versus language-tagged lists - and pin a snapshot date, since the export trails edits to the live record.

Can Datadory deliver both datasets together?

Yes - alone or merged onto one barcode key, delivered daily, weekly, or hourly, your call. Every delivery travels with both verified field dictionaries and sample rows for inspection before anything ships, so you see the exact grain a delivery arrives at. Or take both in one feed.