Steel Data Provider: From Crude-Output Tonnages to the Defect Mask · Head-to-head

Ministry of Steel Monthly Summary vs Severstal Steel Defect Detection

Which steel data provider: from crude-output tonnages to the defect mask data fits your job: Ministry of Steel Monthly Summary, or Severstal Steel Defect Detection. API, files, or your warehouse. Daily, weekly, or hourly.

Steel Data Provider: From Crude-Output Tonnages to the Defect Mask

Ministry of Steel Monthly Summary

Steel Data Provider: From Crude-Output Tonnages to the Defect Mask

Severstal Steel Defect Detection

Where the fields line up

No shared field names. These two answer different questions.

Field Ministry of Steel Monthly Summary Severstal Steel Defect Detection
report_month documented not in this set
crude_steel_production_mt documented not in this set
finished_steel_production_mt documented not in this set
finished_steel_consumption_mt documented not in this set
finished_steel_imports_mt documented not in this set
finished_steel_exports_mt documented not in this set
availability_mt documented not in this set
stock_variation_mt documented not in this set
top_traded_products documented not in this set
country documented not in this set
crude_steel_by_country_mt documented not in this set
cpse_production documented not in this set

Coverage, side by side

Ministry of Steel Monthly Summary Severstal Steel Defect Detection
Granularity One reading per reporting month and indicator - national totals, quarterly fiscal cumulatives, per-country and per-enterprise lines One row per image-class pair - ImageId_ClassId such as '004f40c73.jpg_1', four rows per test image

What each contains

Pick by fit, not by loyalty.

Ministry of Steel Monthly Summary Severstal Steel Defect Detection
Unit of analysis One reading per reporting month and indicator - national totals, quarterly fiscal cumulatives, per-country and per-enterprise lines One row per image-class pair - ImageId_ClassId such as '004f40c73.jpg_1', four rows per test image
Classification scheme Named product categories: HR COIL/STRIP, CR COIL/SHEETS, GP/GC SHEETS/COIL, PLATES, BARS & RODS, OTHERS ClassId 1-4 marking one of four defect types per mask; the semantic meaning of each class was never published
Quantity measured Tonnes: crude steel 42.06 Mt and finished consumption 41.57 Mt for April-June 2026-27 (P); June exports 0.62 Mt against imports 0.70 Mt Pixels: EncodedPixels run-length pairs such as '29102 12 29358 18', counted column-wise over a flattened 256 x 1600 strip
Row identity Country (China, India, United States, Japan, Russia, South Korea, World), CPSE names (SAIL, NMDC, KIOCL, MOIL, RINL, NSL) and report month Hex-named JPG files such as '0002cc93b.jpg'
Completeness convention Recent months flagged provisional (P) until finalized in later editions 1,801 test images whose labels were never released; about 63 percent of training images carry no defect at all

What each does better

the Ministry of Steel Monthly Summary

A whole economy in fourteen fields. Crude steel production of 42.06 Mt and finished production of 40.99 Mt for April-June 2026-27 (provisional), set against consumption of 41.57 Mt - a gap the report closes through imports of 0.70 Mt in June against exports of 0.62 Mt. The fiscal ladder runs readable year over year: finished steel consumption for the same April-June window moved 30.83 to 35.54 to 38.40 to 41.57 million tonnes across four consecutive fiscal years.

World context built in. The comparison annexure sets India's 14.09 Mt of May-2026 crude beside China, the United States, Japan, Russia and South Korea and a world total of 157.88 Mt - down 0.3 percent on May 2025, with 2025 closing at 1,848.9 Mt, down 2.0 percent on 2024.

Named corporate actors. The enterprise annexure reports hot metal, crude steel, saleable steel, pellets and manganese ore for SAIL, NMDC, KIOCL, MOIL, RINL and NSL every month, in lakh tonnes, with cumulative fiscal-year figures beside the monthly ones.

Severstal Steel Defect Detection

Twelve thousand five hundred sixty-eight labeled examples. Every training strip is a 256 x 1600-pixel frame from cameras mounted over a working flat-sheet line, annotated with run-length-encoded masks counted column-wise over the flattened image ('29102 12 29358 18 ...' on a representative row). That is pixel-localized ground truth no statistical table can offer.

The benchmark effect. Backed by a USD 120,000 prize pool funded by Severstal PAO and run from July to October 2019, the competition fixed the evaluation contract the field still uses: mean Dice coefficient computed per ImageId-ClassId pair, defined as 1 when prediction and truth are jointly empty. Surface-defect segmentation baselines, architectures and augmentation recipes still cite it as their reference point.

Imbalance as rehearsal, not noise. About 63 percent of training images contain no defect at all, and the labeled remainder is heavily skewed - class 3 dominant, class 2 rare. Any team building a visual quality gate inherits precisely this distribution in production, so the set trains the failure modes that matter: missing rare defects and over-confident negatives.

The verdict

Verdict: sample both - they tie at 8 out of 10, so the question decides, not the leaderboard.

Pick Severstal Steel Defect Detection when the unit of analysis is the sheet: training and benchmarking segmentation models under Dice, rehearsing a class-imbalanced inspection pipeline before aiming it at a live line, teaching convolutional architectures on real industrial imagery, or reproducing one of the most-cited results in industrial computer vision.

Neither substitutes for the other. The report cannot tell you where a defect sits on a strip; the imagery cannot tell you whether steel demand rose last quarter.

Sample both, pick by fit. See Ministry of Steel Monthly Summary · See Severstal Steel Defect Detection

Fair questions

Is the Ministry of Steel Monthly Summary better than the Severstal Steel Defect Detection dataset?

Different instruments, tied on craft - both score 8 out of 10. The Severstal side wins whenever the question is visual: 12,568 labeled 256 x 1600-pixel strips with run-length-encoded masks across four defect classes. Sample both and match each to the question.

Which dataset is bigger?

It depends which direction you count. Records: the imagery wins, 12,568 training images plus 1,801 unlabeled test strips against roughly eighty-plus monthly report issues. Bytes tilt toward the imagery at about 1.4 GB compressed; information density per record tilts toward the report's economy-wide indicators.

Do the two datasets overlap?

On almost nothing tabular. The report keys rows by reporting month and indicator name; the imagery keys rows by hex filename and defect class. The genuine conceptual overlaps are classification itself - named product categories versus four unpublished defect classes - and incompleteness: provisional months on one side, 1,801 test strips with no released labels on the other. With no shared key, no common specimen and no geographic or temporal intersection, the overlap stays conceptual rather than tabular.

Can I use both datasets in one machine-learning project?

Yes, at the pipeline level rather than the row level. Because the collections share no specimens and no schema, the combination shapes system architecture; it never becomes one merged table.

Can Datadory deliver both datasets together?

Yes. Either record arrives alone or both arrive aligned onto one calendar, delivered daily, weekly, or hourly - your call. Name the indicators, product categories, enterprises and fiscal windows when you request the sample and it lands pre-cut, with field definitions and coverage profiles attached. Or take both in one feed.