Steel Data Provider: From Crude-Output Tonnages to the Defect Mask · Head-to-head

UNSD Industrial Commodity Statistics (steel) vs NEU Metal Surface Defects Database

Which steel data provider: from crude-output tonnages to the defect mask data fits your job: UNSD Industrial Commodity Statistics, or NEU Metal Surface Defects Database. API, files, or your warehouse. Daily, weekly, or hourly.

Steel Data Provider: From Crude-Output Tonnages to the Defect Mask `Country or Area` - roughly 103 reporting geographies for the crude steel series

UNSD Industrial Commodity Statistics (steel)

Steel Data Provider: From Crude-Output Tonnages to the Defect Mask None - imagery collected at Northeastern University (China) with no location attribute recorded

NEU Metal Surface Defects Database

Where the fields line up

No shared field names. These two answer different questions.

Field UNSD Industrial Commodity Statistics NEU Metal Surface Defects Database
country_or_area documented not in this set
year documented not in this set
unit documented not in this set
value documented not in this set
value_footnotes documented not in this set
commodity_series documented not in this set
image file not in this set Grayscale JPEG filename of a hot-rolled steel strip patch, e.g. crazing_1.jpg; 200x200 pixels.
defect class not in this set One of six classes: crazing (Cr), inclusion (In), patches (Pa), pitted surface (PS), rolled-in scale (RS), scratches (Sc); encoded in the folder name and filename prefix.
bounding box (xmin, ymin, xmax, ymax) not in this set Pixel coordinates of each annotated defect region inside the image, stored in per-image PASCAL VOC-style XML files in the ANNOTATIONS folder.
split not in this set In mirrored distributions, train/validation/test assignment; some mirrors provide train_images/test_images folders plus YOLO-format txt labels.

Coverage, side by side

UNSD Industrial Commodity Statistics NEU Metal Surface Defects Database
Geographic `Country or Area` - roughly 103 reporting geographies for the crude steel series None - imagery collected at Northeastern University (China) with no location attribute recorded
Granularity One reading per country or area, year and commodity series - quantity in thousand metric tons and value in millions of USD One row per annotated image - class label parsed from the filename prefix such as crazing_1.jpg, several boxes possible per file

What each contains

Pick by fit, not by loyalty.

UNSD Industrial Commodity Statistics NEU Metal Surface Defects Database
Unit of analysis One reading per country or area, year and commodity series - quantity in thousand metric tons and value in millions of USD One row per annotated image - class label parsed from the filename prefix such as crazing_1.jpg, several boxes possible per file
Classification scheme `Country or Area` plus placement on the UN List of Industrial Products: crude steel, hot- and cold-rolled flat products, bars and rods, wire, tubes, coated sheet, electrical steel `defect class` enum: crazing (Cr), inclusion (In), patches (Pa), pitted surface (PS), rolled-in scale (RS), scratches (Sc)
Measured quantity `Value` - the production observation itself, in whichever `Unit` the row declares `bounding box (xmin, ymin, xmax, ymax)` - pixel coordinates of each defect region inside a 200x200 grayscale frame
Period stamp `Year` - reference year, 1995-2016 in the online database, with 1950-2003 history on request None - the set was assembled circa 2010s and frozen; no acquisition date rides on any image
Geography `Country or Area` - roughly 103 reporting geographies for the crude steel series None - imagery collected at Northeastern University (China) with no location attribute recorded
Provenance / caveat flag `Value Footnotes` - codes marking estimated, breakdown or different-source values `split` - train/validation/test assignment present in mirrored distributions
Row identity Composite of geography, year and product series `image file` - the JPEG filename, e.g. inclusion_23.jpg
Modality Tabulated numeric series delivered as typed columns Grayscale JPEGs paired with PASCAL VOC-style XML annotations

What each does better

UNSD Industrial Commodity Statistics

Cross-country economics no image set can serve. The crude steel series spans roughly 103 reporting countries over 22 online years, and each observation arrives twice over - physical quantity in thousand metric tons plus monetary value in millions of USD - so one table answers both how much a country made and roughly what it was worth. Capacity studies, market sizing and demand models all read off the same rows.

Product resolution beyond crude steel. Dozens of companion series decompose the industry downstream: hot-rolled and cold-rolled flat products, bars and rods, angles, shapes and sections, sheet piling, railway track material, wire of iron or non-alloy steel, tubes and pipes, tin- lead- and zinc-plated galvanized sheets, silicon-electrical steel. A supply-chain question can follow a specific product family across all of them.

Honesty flags on the numbers. Every value carries a Value Footnotes code qualifying estimates, breakdowns and differing sources - the difference between quotable and caveated in any brief.

Institutional pedigree. Compiled by the United Nations Statistics Division from national statistical offices, with historical series reaching back to 1950 available on request - a longer memory than most national releases publish.

the NEU Metal Surface Defects Database

Ground truth pixels, not proxies. 1,800 grayscale 200x200 images of real hot-rolled strip, each paired with PASCAL VOC-style XML annotation storing xmin, ymin, xmax, ymax for every defect region - multiple instances per image supported, which is what makes detection work possible alongside plain classification. No tabular release can train a vision model.

Benchmark comparability. This is the set industrial anomaly- and defect-detection papers standardize on through the NEU-DET variant, so a score measured on it lands next to years of published results rather than floating alone. Balanced design helps too: 300 samples per class takes class imbalance out of the experimental plan before it starts.

Honest difficulty. The six chosen classes are the ones that resist easy separation under mill conditions - low-contrast crazing, texture-like patches, thin directional scratches - which is precisely why a strong result here means something on a real line.

Teaching weight. At about 60 MB compressed the entire benchmark fits a laptop, and a complete exercise - data, labels, trained model - fits inside one course module.

The verdict

Verdict: sample both - the question decides, not the leaderboard, and these two never enter the same race.

Pick UNSD Industrial Commodity Statistics (steel) when the unit of analysis is the market: capacity and output comparisons across countries, import-substitution and demand studies, product-mix tracking from crude through galvanized and electrical steel, or any model that needs a country-year panel of production in tonnes and dollars.

Pick NEU Metal Surface Defects Database when the unit of analysis is the material: training and benchmarking classification or detection models on strip-surface defects, rehearsing an inspection pipeline before aiming it at a live line, evaluating labeling tools against fixed ground truth, or reproducing one of the most-cited results in industrial computer vision.

Three quick tests settle most cases. Does your question name a country or a year? Only the statistics travel. Does it name crazing or scratches? Only the imagery knows what those look like. Need to size the market for inspection automation and then prove the concept on real defect imagery? That is the both-of-them case, and it comes up constantly.

Sample both, pick by fit. See UNSD Industrial Commodity Statistics · See NEU Metal Surface Defects Database

Fair questions

Is UNSD Industrial Commodity Statistics (steel) better than the NEU Metal Surface Defects Database?

The UNSD record wins whenever the question is economic: annual country-level production of crude steel and rolled products in thousand metric tons and millions of USD across roughly 103 countries from 1995 to 2016. The NEU record wins whenever the question is visual: 1,800 grayscale 200x200 strip images across six defect classes with per-instance bounding boxes. Sample both and match each to its question.

Do UNSD Industrial Commodity Statistics (steel) and the NEU Metal Surface Defects Database overlap?

On almost nothing. The UNSD dictionary keys rows by country or area, year, unit and value; the NEU dictionary keys rows by image filename, defect class, bounding box and split assignment. No shared identifier exists, no country appears in the imagery and no tonnage appears in the annotations. What they share is an industry and a habit of classification - product families on one side, six named defect classes on the other - so the overlap stays conceptual rather than tabular.

Which dataset is bigger?

It depends which direction you count. Records: the crude steel series alone holds 2,300 observations, with dozens of further product series carrying hundreds to roughly 2,000 each, against 1,800 images. Columns: five documented fields against four. Bytes tilt toward the imagery at about 60 MB compressed; information breadth per record tilts toward the statistics, where one row carries a country's entire year of output in quantity and value.

Can I use both datasets in one project?

Yes, as two layers of one story rather than one joined table, because no key travels between them. A defensible pairing sizes the market first - which countries produce how much crude steel, at what value - then justifies an inspection investment using the six-class benchmark as published evidence of what machine vision catches on hot-rolled strip. Both records are closed-ended, so pair them with fresher complements before citing either against current conditions.

Can Datadory deliver both datasets together?

Yes. Either record arrives alone or both arrive aligned onto one calendar, delivered daily, weekly, or hourly - your call. Name the countries, product series, defect classes and year windows when you request the sample and it lands pre-cut, with field definitions and coverage profiles attached. Or take both in one feed.