Construction Materials Data Provider · Head-to-head
27-Year Multi-Producer European Cement Characterization Dataset vs Concrete Compressive Strength Dataset
Which construction materials data provider data fits your job: 27-Year Multi-Producer European Cement Characterization Dataset, or Concrete Compressive Strength Dataset. API, files, or your warehouse. Daily, weekly, or hourly.
27-Year Multi-Producer European Cement Characterization Dataset
Concrete Compressive Strength Dataset
What each contains
Pick by fit, not by loyalty.
| 27-Year Multi-Producer European Cement Characterization Dataset | Concrete Compressive Strength Dataset | |
|---|---|---|
| Compressive strength | `compressive_2d`, `compressive_7d`, `compressive_28d` - three checkpoints per sample; one 52.5-class CEM I reads 30.5 / 46.3 / 58.7 MPa | `Concrete compressive strength(MPa)` - one measured value per row; a 540 kg/m3 cement mix reads 79.99 MPa at 28 days |
| Binder description | Oxide suite stored as mass fractions - `chem_CaO` 0.625, `chem_SiO2` 0.205 - plus loss on ignition, drying loss and the minor-oxide tail | `Cement (component 1)(kg in a m^3 mixture)` at 540.0, beside blast furnace slag 142.5 and fly ash dosages |
| Mineralogy | Corrected Bogue phases (`c3s` 0.4001, with C2S, C3A, C4AF) plus lime saturation factor 0.8614 and the silica and alumina moduli | None - slag and fly ash appear only as quantities, never as phases |
| Water | `water_demand` at standard consistency, stored as a unit fraction (0.266) | `Water (component 4)(kg in a m^3 mixture)` dosed per batch (162.0 to 228.0 across sample rows) |
| Fineness and particles | `blaine` 4140 cm2/g, `x50` median 6.01 um, Rosin-Rammler location and uniformity parameters | None - aggregates enter as coarse and fine masses (1040.0 / 676.0) with no size descriptors |
| Time axis | `delivery_date` calendar stamps spanning 1997-2024, strengths read at fixed 2/7/28-day ages | `Age (day)` curing age from 1 to 365, with no calendar date anywhere |
| Row identity | `company` (Company_1), `cement_type` (CEM II/A-LL), `strength_class` (52.5) and an `is_cem1` flag | — |
| Admixtures | Not represented among the 50 variables | `Superplasticizer (component 5)(kg in a m^3 mixture)` dosed explicitly (2.5 on sample rows) |
What each does better
the 27-Year Multi-Producer European Cement Characterization Dataset
Fifty variables against nine. Every sample ships the full characterization stack: major oxides (CaO 0.625 and SiO2 0.205 on a typical row), the minor-oxide tail down to TiO2, MnO, MgO, chlorides and uncombined lime, loss on ignition, density, Blaine surface area, particle-size descriptors through Rosin-Rammler location and uniformity, initial and final setting times, standard-consistency water demand, Le Chatelier soundness - and then a derived layer nobody has to recompute: lime saturation factor, silica and alumina moduli, equivalent alkali and corrected Bogue phase fractions (C3S 0.4001 on that same row). The concrete table offers eight dosages and an age.
Twenty-eight years of industrial practice, held constant. Deliveries run from 1997 to 2024 across 23 producers whose labels stay stable while their names stay hidden - Company_1 in 2001 is Company_1 in 2018. That is the design drift studies need: how Blaine creeps, how setting behaviour shifts, how one producer's 52.5-class cement lands at 58.7 MPa in one year and 63.4 MPa a few years later. Around the 311-sample CEM I backbone sit 165 CEM II, CEM III and specialty rows.
Strength explained, not just predicted. Because chemistry and mineralogy ride along with every strength result, the cement table supports why-questions: what separates the strong sample from its weaker sibling, what a setting-time ratio does to early-age gain, how equivalent alkali tracks the clinker behind it. Nine mix quantities cannot ask those questions.
the Concrete Compressive Strength Dataset
More experiments, wider spread. 1,030 mixes against 476 samples - and the concrete collection varies its inputs deliberately: cement from lean to rich, slag and fly ash swapped in and out, superplasticizer present or absent, water and aggregates moved around. For fitting a regression, which is precisely the job this collection was assembled for, breadth across the input space beats depth per specimen.
Continuous age. The cement table reads strength at 2, 7 and 28 days, full stop. The mix table records curing ages from 1 to 365 days, so strength-gain curves are fittable rather than interpolated. Three sample rows show the nonlinearity the repository itself warns about: 540 kg/m3 of cement at 28 days reads 79.99 MPa; 332.5 kg/m3 with slag at 270 days reads 40.27; 198.6 kg/m3 at 360 days manages 44.3. Age is doing as much work as composition.
Nine columns, zero gaps. No missing values, no reshaping, no unit traps beyond reading the headers - the whole thing fits in memory and trains a baseline in seconds. As a teaching instrument or a regression-test fixture it stays the classic for a reason: every reviewer in materials machine learning has seen it before.
The verdict
Verdict: sample both, pick by fit - they tie at 8 out of 10, so the question decides, not the leaderboard.
Pick the 27-Year Multi-Producer European Cement Characterization Dataset when the unit of analysis is the binder: explaining why one 52.5-class sample reached 63.4 MPa while another stopped at 58.7, modelling grind and mineralogy effects, auditing supplier consistency across decades, or building interpretable binder features for any downstream strength model.
Pick the Concrete Compressive Strength Dataset when the unit of analysis is the mix: predicting strength from proportions, fitting strength-age curves across a full year of curing, benchmarking a fresh regression method on a dataset every reviewer already knows, or teaching the nonlinearity of age and composition in one lecture.
Three quick tests settle most cases. Need oxide chemistry or Bogue phases? Only the cement side has them. Need curing ages past 28 days? Only the mix side has them. Need explanation and prediction in the same project? That is the both-of-them case.
Sample both, pick by fit. See 27-Year Multi-Producer European Cement Characterization Dataset · See Concrete Compressive Strength Dataset
Fair questions
Is the 27-Year Multi-Producer European Cement Characterization Dataset better than the Concrete Compressive Strength Dataset?
Different instruments, tied on craft - both score 8 out of 10. The cement set wins whenever the question is explanatory: 50 chemical, physical and mineralogical variables on 476 production samples spanning 1997 to 2024. The concrete set wins whenever the question is predictive scale: 1,030 mixes, nine attributes, no missing values and curing ages from 1 to 365 days. Sample both and match each to the question.
Which dataset is bigger?
It depends which direction you count. Rows: the concrete set, 1,030 against 476. Cells: the cement set, roughly 23,800 values against 9,270. Information per row tilts heavily toward the cement table's 50 variables, while most of the concrete table's rows earn their place as genuinely distinct experimental points.
Do the two datasets overlap?
On the target quantity - compressive strength in MPa - and little else. The cement side measures it at 2, 7 and 28 days per sample; the concrete side reports one value per mix at an age between 1 and 365 days. With no shared key, no common specimen and no geographic or temporal intersection, the overlap stays conceptual rather than tabular.
Can I use both datasets in one machine-learning project?
Yes, at the model level rather than the row level. A common pattern fits mix-proportion regressions on the concrete records, then uses the cement table's chemistry and Bogue-phase features to explain residual structure or design better binder-side features. Because the collections share no specimens, the combination shapes modelling choices; it never becomes one merged table.
Can Datadory deliver both datasets together?
Yes. Either record arrives alone or both arrive merged onto one calendar, delivered daily, weekly, or hourly - your call. Name the producers, mix designs and year windows when you request the sample and it lands pre-cut, with field definitions and coverage profiles attached. Or take both in one feed.