27-Year Multi-Producer European Cement Characterization Dataset

Datadory delivers 27 year multi producer european cement characterization dataset data: 476 laboratory-tested cement samples from 23 anonymized European producers spanning 1997 to 2024. Each sample row carries oxide chemistry, Blaine fineness, setting behaviour, derived clinker phases and 2-, 7- and 28-day strength results across 50 documented variables.

Where it covers
Europe - 23 anonymized cement producers (countries withheld)
How far back
Laboratory measurements 1997-2024
How fine
One row per cement sample - 476 samples, 50 variables

What is the 27-Year Multi-Producer European Cement Characterization Dataset?

Almost every cement laboratory archive dies in a filing cabinet: thousands of routine quality-control measurements, taken every shift for decades, useful to nobody outside the plant that took them. This dataset is what happens when somebody rescues one. Researchers at Bauhaus-Universitat Weimar and the Technical University of Munich assembled 476 cement samples from 23 anonymized European producers, covering laboratory deliveries from 1997 through 2024, into a single processed table of 50 variables per sample.

The split is 311 CEM I samples against 165 other records - mostly CEM II and CEM III, with a small number of CEM IV, CEM VI and special cements. Each row binds together what the lab actually measured: declared type and strength class, full oxide chemistry, density and Blaine fineness, particle-size descriptors where measured, setting times, water demand, Le Chatelier soundness, and flexural and compressive strength at 2, 7 and 28 days. Derived engineering features ride alongside - lime saturation factor, silica and alumina moduli, equivalent alkali and corrected Bogue phase fractions - so the table supports clinker-modeling work without re-deriving anything.

The multi-producer, multi-decade design is the rare part. Single-factory archives cannot tell you whether a chemistry-to-strength relationship generalizes across plants; this one can. Get a sample of this dataset and pull the CEM I rows first - they are the backbone the accompanying machine-learning study was built on.

What does a sample row look like?

One row per cement sample, flat tabular, no parsing gymnastics. Three illustrative rows from the same anonymized producer:

company        : Company_1         cement_type    : CEM I
strength_class : 52.5              delivery_date  : 2001-05-28
density        : 3.132 g/cm3       blaine         : 4140 cm2/g
chem_CaO       : 0.625             chem_SiO2      : 0.205
compressive_28d: 58.7 MPa          is_cem1        : True

company        : Company_1         cement_type    : CEM I
strength_class : 52.5              delivery_date  : 2004-03-12
density        : 3.118 g/cm3       blaine         : 4230 cm2/g
chem_CaO       : 0.629             chem_SiO2      : 0.201
compressive_28d: 63.4 MPa          is_cem1        : True

company        : Company_1         cement_type    : CEM II/A-LL
strength_class : 52.5              delivery_date  : 2018-10-16
density        : 3.095 g/cm3       blaine         : 6080 cm2/g
chem_CaO       : 0.632             chem_SiO2      : 0.197
compressive_28d: 61.6 MPa          x50            : 6.01 um

Values above quote the delivered table directly. Read the three rows together and the longitudinal design shows itself: within one anonymous producer, Blaine fineness climbs from 4140 to 6080 cm2/g between a 2001 CEM I and a 2018 CEM II/A-LL while calcium oxide holds near 0.63 and 28-day strength stays in a narrow 59-63 MPa band. That is seventeen years of grinding technology and binder evolution captured in comparable units - the kind of panel a single-year snapshot simply cannot give you. Chemistry fractions are stored on a 0-1 scale, with weight-percent documentation carried in the shipped data dictionary.

What fields does each record include?

Fifty variables make up the full schema; the twenty-four columns below are the verified spine covering identity, chemistry, physicality, setting behaviour and mechanical performance. Every field carries its unit, storage scale and derivation notes in the publisher's own data dictionary, which travels with the table. Columns beyond the spine - the minor oxide assays, the remaining particle-size percentiles, the rest of the Bogue set - are itemized in the final row.

What does coverage look like across geography, time and granularity?

Geography - Europe, sampled across 23 cement producers. Producer identities are anonymized by design and the countries involved are deliberately not disclosed, so treat this as a continental panel rather than a country-mapped one. If your question needs national attribution, pair it with asset-level sources such as the global cement and concrete tracker.

Temporal - laboratory deliveries from 1997 to 2024. Twenty-eight years is long enough to span several generations of grinding technology, clinker chemistry adjustment and the industry's gradual shift toward blended CEM II and CEM III binders, which means the table supports genuine trend work rather than a single-period correlation exercise.

Granularity - one row per cement sample: 476 rows, 50 variables, 311 of them CEM I. Sample-level resolution is what makes the strength-modelling use case work; aggregate tables average away exactly the variance a regression needs.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Who uses this data, and for what?

  • Machine-learning strength prediction - train and validate models that infer 28-day compressive strength from early-age results and routine chemistry. The publishers themselves framed the dataset as a probe of how far routine cement characterization can be pushed, which makes it a ready-made benchmark; methodology notes live on our data scientists use cases page.
  • Binder evolution research - track how CEM I versus blended cements drifted across 1997-2024 in fineness, sulfate balance and setting behaviour, with producer held constant via the anonymized labels.
  • Standards and specification work - examine how declared strength classes map onto measured 2, 7 and 28-day performance across many producers at once, the cross-plant view single-lab studies lack.
  • Mix design and constructability studies - standard-consistency water demand, initial and final setting times and soundness arrive per sample, so rheology-relevant inputs need no separate sourcing.
  • Decarbonization assessments - the 165 non-CEM-I rows give clinker-substituted binders the same measured treatment as portland cement, useful baseline material for lower-carbon binder performance claims.

Which personas get the most value?

Data scientists and ML engineers get a cleaned, dictionary-documented, sample-level table with derived features already computed - the tedious half of cement modelling done before you open a notebook. Materials R&D and quality engineers get a cross-producer reference distribution for every routine lab measurement, something no individual plant's archive can provide. Developers and data-product builders get a compact, stable schema that drops straight into feature stores and dashboards. Journalists, academics and students get citable laboratory evidence behind claims about European cement quality trends; more angles on our journalists academics students page.

How does it compare to other construction materials datasets?

Most of this industry's data measures tonnes, euros or plants: production series, price trackers, asset registries. This dataset measures the material itself, at laboratory resolution, per sample. The nearest neighbor is the classic 1,030-row Concrete Compressive Strength Dataset - we run the comparison in detail on our vs Concrete Compressive Strength Dataset page; the short version is real multi-producer production samples versus a smaller, coarser mix-proportion table.

Used together they bracket the problem: production and asset sources say who makes cement and how much; this table says what the product actually tests like. The full slate sits on our best construction materials datasets ranking.

What should I know before requesting a sample?

Three things. First, anonymity cuts both ways: producer labels let you hold a company constant across decades, but no country or region attribute exists anywhere in the table, so geographic segmentation beyond 'Europe' is off the table.

Second, expect structural zeros rather than dirt: particle-size descriptors are populated only where PSD was measured, and the shipped data dictionary documents every unit, storage scale and missing-value convention - read it before writing imputation logic, because fraction-valued columns stored 0-1 will silently mislead anyone assuming weight percent.

Third, this is a longitudinal archive through 2024 deliveries, not a rolling series - plan analyses around its fixed historical window and let scheduled delivery handle freshness of your copy.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - 24 verified core columns of the 50-variable schema, one row per cement sample
fieldtypedefinitionexample
companystringAnonymized producer label assigned before analysis.Company_1
cement_typestringDeclared cement type or family per EN 197.CEM II/A-LL
strength_classnumberDeclared 28-day strength class in MPa.52.5
delivery_datedateDate associated with the sample delivery or laboratory record.2001-05-28
densitynumberCement density in g/cm3.3.132
blainenumberBlaine specific surface area in cm2/g.4140
setting_time_startnumberInitial setting time in hours.2.0
setting_time_endnumberFinal setting time in hours.2.42
water_demandnumberStandard-consistency water demand as a unit fraction.0.266
volume_stabilitynumberLe Chatelier soundness in mm.1.75
compressive_2dnumber2-day compressive strength in MPa.30.5
compressive_7dnumber7-day compressive strength in MPa.46.3
compressive_28dnumber28-day compressive strength in MPa.58.7
tensile_28dnumber28-day flexural strength in MPa.9.3
LOInumberLoss on ignition as a mass fraction stored 0-1 (documented as wt%).0.027
LSFnumberLime saturation factor, a clinker modulus computed from the major oxides.0.8614
c3snumberCorrected alite (C3S) Bogue phase fraction, stored 0-1.0.4001
chem_CaOnumberCalcium oxide content as a mass fraction stored 0-1.0.625
chem_SiO2numberSilicon dioxide content as a mass fraction stored 0-1.0.205
sodium_equivalencenumberEquivalent alkali content combining sodium and potassium oxides.0.006261
is_cem1booleanIndicator flag for CEM I samples.True
setting_time_rationumberDerived ratio of final to initial setting time.1.2083
x50numberMedian particle size in micrometres; empty where PSD was not measured.6.01
xfreqnumberRosin-Rammler uniformity parameter, dimensionless; populated only where PSD was measured.--
additional fields on requestvariesRemaining columns among the 50 variables: minor oxide assays (Al2O3, Fe2O3, SO3, K2O, Na2O, water-soluble alkalis, TiO2, MnO, MgO, Cl, uncombined lime), drying loss, further particle-size descriptors (mean, modal, d10, d90, Rosin-Rammler location), 2- and 7-day flexural strengths, the remaining corrected Bogue phases (C2S, C3A, C4AF) and the silica and alumina moduli. Specify the columns you need when you request a sample.per-request

Questions buyers ask

How many producers and samples does the dataset contain?

476 cement samples from 23 European producers, gathered from routine laboratory measurements taken between 1997 and 2024. Each sample is one row with 50 variables, giving roughly a thousand producer-years of compressed laboratory history in a single table.

Are producer identities really hidden?

Yes. Producers appear only as anonymized labels such as Company_1, assigned before analysis, and internal cement codes were removed as well. The label is nonetheless stable per producer, so you can hold a company constant across decades without ever learning which company it is.

Which cement types are covered?

311 of the 476 samples are CEM I. The remaining 165 are primarily CEM II and CEM III blends, with a small number of CEM IV, CEM VI and special cements. Declared type and declared 28-day strength class are recorded per sample alongside the measured values.

What are the corrected Bogue phase fractions?

Bogue phases - C3S alite, C2S belite, C3A and C4AF - are potential mineral phases calculated from oxide chemistry rather than measured directly. The dataset ships corrected variants of that calculation, so clinker-mineral modeling starts from the authors' adjusted figures instead of the textbook formula.

Why are some particle-size cells empty?

Particle-size distribution was measured on a subset of samples, so descriptors such as the median grain size and the Rosin-Rammler parameters are blank where the test was not performed. The data dictionary documents the convention explicitly, which keeps missingness interpretable rather than accidental.

Can I predict 28-day strength from earlier-age results?

That is precisely the modelling question the dataset was published alongside. With 2-, 7- and 28-day compressive strengths plus setting times and chemistry per sample, you can test early-age inference yourself - and the accompanying machine-learning study documents how far that inference realistically reaches.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing