Diversified Chemicals Data Provider · Head-to-head
VCI - German Chemical Industry statistics (Chemiewirtschaft in Zahlen) vs BASF-AI Chemical Data collection (Hugging Face)
Which diversified chemicals data provider data fits your job: VCI - German Chemical Industry statistics, or BASF-AI Chemical Data collection. API, files, or your warehouse. Daily, weekly, or hourly.
VCI - German Chemical Industry statistics (Chemiewirtschaft in Zahlen)
BASF-AI Chemical Data collection (Hugging Face)
Where the fields line up
No shared field names. These two answer different questions.
| Field | VCI - German Chemical Industry statistics | BASF-AI Chemical Data collection |
|---|---|---|
Produktion (production by segment) | documented | not in this set |
Ausgewaehlte Produktionszahlen | documented | not in this set |
Produktionsindizes ausgewaehlter Industriezweige | documented | not in this set |
Index der Erzeugerpreise | documented | not in this set |
Index der Exportpreise / Importpreise | documented | not in this set |
Energie- und Rohstoffeinsatz | documented | not in this set |
Aussenhandel | documented | not in this set |
Beschaeftigung und Einkommen | documented | not in this set |
Umsatz | documented | not in this set |
Forschung und Entwicklung | documented | not in this set |
Investitionen (chapter 07) | documented | not in this set |
Finanzdaten (chapter 09) | documented | not in this set |
What each contains
Pick by fit, not by loyalty.
| VCI - German Chemical Industry statistics | BASF-AI Chemical Data collection | |
|---|---|---|
| Documented fields | 10 | 11 |
| Field types | All numeric: values, volumes in tonnes, index points, billions of euros | — |
| Signature fields | Ausgewaehlte Produktionszahlen (product output); Index der Erzeugerpreise; Aussenhandel; Umsatz; Forschung und Entwicklung | SMILES / InChI / IUPACName; CID; doi; abstract; paragraph |
| Shared concepts | Segment taxonomy (polymers, inorganics, coatings, fibres, fine/specialty) as row labels | None - compound and paper records carry no economic attributes |
What each does better
Chemiewirtschaft in Zahlen
The BASF-AI collection contains no euro, no tonne and no employee anywhere in its schemas.
It has a time dimension the ML corpora lack. Continuous multi-year series per chapter, 'Stand'-dated December 2025 through August 2026, with archived editions back to 1955 for longitudinal work. The collection's temporal story is a snapshot: last collection update June 25, 2025, individual datasets touched between April 2025 and January 2026.
It distinguishes two statistical boundaries. Six worksheets in the production chapter separate the VCI boundary from the official Destatis one - a reconciliation nobody else publishes, useful whenever your German chemicals figures disagree with the federal statistics office's.
It carries product-level resolution inside its aggregate frame. Tonnes for chlorine, methanol, vinyl chloride, acetic acid, ethylene dichloride, ethylene oxide, ethylene glycol and propylene glycol/oxide, plus production indices benchmarking chemicals against Maschinenbau and other sectors.
the BASF-AI collection
Record-level scale: millions of rows versus a few hundred indicators. PubChem-Raw alone carries 2,087,164 compounds with CID, Title, MolecularFormula, IUPACName, InChI, SMILES and Synonyms, plus 408,530 descriptions wired to SourceName, SourceID and ReferenceNumber provenance. The whole collection totals several gigabytes. Chemiewirtschaft in Zahlen publishes industry totals - there is no molecule to look up in it.
Structural chemistry as first-class fields. InChI and SMILES strings on every compound row make substructure search, fingerprinting and model featurization direct column operations. The VCI workbooks have no analogue; their atoms are euros and index points.
It ships ready-made ML training material, not just raw text. dolma-chemistry-only contributes 1,187,726 paragraph records; ChemRxivRetrieval is a full BEIR-style evaluation harness - 69,457 passages, 5,000 queries, query-corpus relevance scores - so a retrieval system can be benchmarked the day you load it. Nothing on the statistics side trains anything.
Machine-native packaging. Parquet throughout, with per-dataset README dataset_info YAML documenting feature names, dtypes and split sizes - loadable straight into pandas or the HF datasets library. The German side is human-first OOXML: multi-sheet workbooks with shared strings, German headers and 'Stand' dates that reward careful parsing.
Where they're equivalent
Both come from industry insiders. VCI is the lobby and statistics body for Germany's roughly 2,000 chemical-pharmaceutical companies; BASF-AI is the research arm of the world's largest chemicals producer. Neither is a government office and neither is a data vendor.
Both are verified catalog entries, field definitions confirmed rather than inferred, scoring 8/10 and 7/10 on Datadory's rubric against a catalog-wide average near 7.8.
Both are chemistry-adjacent but not overlapping - which is why this page exists. Their field dictionaries intersect only in the loosest sense: a compound name like Aspirin appears in PubChem titles and synonyms, while the same substance sits somewhere inside VCI's fine and specialty chemicals segment totals. No field in one answers a question in the other.
Neither carries person-level or company-level microdata. VCI aggregates to segments, products and the industry; BASF-AI aggregates to compounds, papers and paragraphs. Both stop short of naming firms.
The verdict
Verdict: sample both, pick by fit - they are different instruments pointed at the same industry.
Accept German-first labeling, workbook archaeology across twelve files, and indicator-level rather than record-level granularity.
Take the BASF-AI collection if you build things that read chemistry. Domain-adaptive pretraining on Dolma-derived corpora, retrieval or embedding benchmarks over preprint passages, cheminformatics pipelines keyed on SMILES and InChI, or LLM tooling grounded in compound description pairs. Accept non-commercial terms on most tagged datasets, a handful of untagged ones needing confirmation, and no economic dimension whatsoever.
If your question starts with 'how big', 'how expensive', 'how many jobs' or 'since when' - the VCI statistics. If it starts with 'what does this molecule look like' or 'can my model read a patent' - BASF-AI.
Sample both, pick by fit. See VCI - German Chemical Industry statistics · See BASF-AI Chemical Data collection
Or take both in one feed
Yes - sequentially rather than side by side, because they occupy different stages of the same workflow.
Run the loop backwards for R&D screening: mine dolma-chemistry-only paragraphs for emerging topics, then check whether VCI's production indices show industrial uptake.
Two cautions from the records. First, join keys do not exist between them: VCI reports segment and product labels in German, BASF-AI reports identifiers, so any bridge is a mapping you maintain. Second, cadences differ by design - continuously refreshed 'Stand' dates through August 2026 on one side, a static June 2025 collection snapshot on the other - so treat the molecular layer as a fixed reference set, not a moving series.
Or take both in one feed: Datadory delivers them alongside the rest of the diversified chemicals catalog, normalized to one schema, on the schedule your team sets.
API, files, or your warehouse. Daily, weekly, or hourly.
Fair questions
Is Chemiewirtschaft in Zahlen better than the BASF-AI collection?
Better for different questions entirely. The BASF-AI collection wins when you need chemistry itself computable: 2,087,164 compounds with SMILES and IUPAC names plus 30,378 ChemRxiv preprints. One has no molecules; the other has no euros.
Do Chemiewirtschaft in Zahlen and the BASF-AI collection cover the same ground?
No, and that is the point of comparing them. Their field dictionaries barely intersect: VCI publishes numeric industry indicators under German labels while BASF-AI publishes compound identifiers, structural strings and paper text. A substance such as aspirin sits inside a VCI fine-chemicals subtotal and also owns a full PubChem record - the same fact at two zoom levels, never in the same columns.
Which one is bigger?
The BASF-AI collection, by orders of magnitude in rows: several GB across ten datasets including 2,087,164 PubChem compound rows and 1,187,726 corpus paragraphs. Chemiewirtschaft in Zahlen wins on depth in time instead - twelve workbooks of continuously refreshed series plus archived annual editions reaching back to 1955.
Which one changes more often?
The VCI statistics. Chapters carry 'Stand' dates from December 2025 to August 2026 and refresh continuously through the year. The BASF-AI collection was last updated June 25, 2025, with constituent datasets touched between April 2025 and January 2026 - a reference snapshot rather than a moving series. Either can be delivered daily, weekly, or hourly.
Can I get both Chemiewirtschaft in Zahlen and the BASF-AI collection from Datadory?
Yes - sample both and pick by fit, since they solve different problems, or take both in one feed. Datadory normalizes each to its documented field dictionary, attaches sample rows for validation, and ships them on the delivery schedule your workflow needs, next to the rest of the diversified chemicals catalog.