Datadory notebook
World Mining Data Excel download: what the workbooks hold, delivered clean
1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.
1,744 datasets. Pick your catch.
What do the delivered rows look like?
Rows exactly as they arrive:
# one observation = one country x commodity x year
dataset : wmd_production_by_country
commodity_sheet: Copper
country : Chile
unit : metr. t
y_2020 : 5733100
y_2021 : 5624900
y_2022 : 5330400
y_2023 : 5250400
y_2024 : 5505890 # r = reported
commodity_sheet: Copper
country : China
unit : metr. t
y_2020 : 1942000
y_2021 : 1964400
y_2022 : 1925600
y_2023 : 1830000
y_2024 : 1842900 # r
# second record type: market-share ranking per commodity
record_type : share_of_world_production_copper_2024
rank_2024 : 1
rank_2023 : (1)
country : Chile
unit : metr. t
production_2024: 5505890
share_pct : 24.03
hhi_contribution: 577.51
rank_2024 : 2
country : Congo, D.R.
production_2024: 3100234
share_pct : 13.53
rank_2024 : 3
country : Peru
production_2024: 2736237
share_pct : 11.94Illustrative shape - request a sample and verified current-edition values come back cut to the commodities and countries you name.
Three properties in these rows decide whether the series integrates quietly or fights back. First, the unit travels with the value - copper states metr. t while gold and platinum-group metals state kg, diamonds state ct and natural gas states Mio m3 - so no join ever assumes tonnes everywhere. Second, the quality flag sits in-band: r reported, e estimated, p provisional, which is what stops a dashboard charting a provisional figure as if it were settled. Third, both record types key on the same commodity-and-country pair, so attaching a world share or a concentration contribution to a production series is a join, not an afternoon.
Which fields does the production panel carry?
Twelve core fields cover both record types - the country-by-commodity production rows and the per-commodity ranking rows - and the ones doing most analytical weight appear below. Each underlying worksheet self-describes its layout in its leading cell, so column meaning survives translation into delivered form intact. Fields beyond the core twelve - price-series context, finer geographic aggregations, longer historical spans - ride along as additional fields on request once scoped, joining the same feed rather than spawning a second pipeline. The full dictionary with examples ships attached to every sample.
How deep does world mining data go - and where is the seam?
Temporally there are two depths, and mistaking one for the other is the classic error in this dataset.
- World totals by continent run from 1984 through the reference year (2024 in the current edition) - the long series, and the only one that supports four-decade narrative charts.
- Country-by-commodity detail covers the five most recent years, 2020-2024. A forty-year country panel does not exist in any single vintage. Building one means harvesting older editions and reconciling restatements against the r/e/p flags - legitimate decade-scale work, and scoped as an additional field set on the same feed when a project needs it.
Geographically the coverage is genuinely global, and the rollups are part of the product rather than something you construct: each of the 168 producer countries rolls up to continent, world-region code, development status, income band, political-stability class and economic bloc (BRICS, EC, OECD among them). "Non-ferrous output by income class" or "mineral fuel production inside versus outside the EC" answers from the same cube as plain country totals.
One seam runs through any multi-year stitch: every edition revises the full table set rather than appending to it, so last cycle's workbook is superseded, not extended. Pin the vintage your model assumes and stamp it beside the values. Handled as delivered rows, the feed carries the newest vintage forward automatically and labels provenance per observation, so archived snapshots stay dated instead of silently stale.
Reading rules that save rework
- Recoverable content, not ore. Quantities state the contained mineral after processing, not rock moved. Comparing them against an ore-based series without converting makes every number look wrong by an order of magnitude.
- The basis is written into the name. Where chemistry differs from elemental tonnes, the commodity label says so: Chromium (Cr2O3), Lithium (Li2O), Rare Earths (REO), Potash (K2O), Uranium (U3O8). Read the label before quoting a figure in anything a client will see.
- Units change per commodity. Metr. t for copper, kg for gold and platinum-group metals, ct for diamonds, Mio m3 for natural gas - the unit rides beside every delivered value precisely so nobody aggregates across them by accident.
- Trust the flag, not the hope. Reported (
r) figures are safe to anchor on; estimated (e) and provisional (p) values get revised between editions, so label them in any dashboard and expect movement. - Rankings double as checksums. Shares descend monotonically down the ranking and HHI contributions sum toward the commodity's concentration index; if a rebuilt extract breaks either property, the unpivot lost a row somewhere upstream.
A sample request returns these populated with real quoted years rather than placeholder cells, cut to whichever commodities and countries your model needs.
How is world mining data delivered through Datadory?
API, files, or your warehouse. Daily, weekly, or hourly.
Spreadsheet-native records are the norm here, not the exception - roughly 410 of the 1,744 datasets in Datadory's catalog ship spreadsheets against only 86 shipping Parquet - so format conversion happens upstream of your warehouse rather than inside it. The identical typed rows land as XLSX for a spreadsheet-bound memo, CSV for a quick load, Parquet for a training pipeline, or warehouse-native writes for the dashboard, and changing format or frequency later is configuration, not migration.
Because production figures sit under continuous revision between editions, the latest vintage always wins - the feed carries it forward so nothing anchors on superseded numbers, and any snapshots you keep arrive stamped with their own dates.
Which records carry world mining data side by side?
Six records bracket the ground a mining production question walks, and the choice among them decides the grain, the geography and the companion measures. The table lines up what each contributes - read it as a routing decision, not a ranking. When the question narrows to United States mine-level safety and inspection events rather than national output, MSHA vs World Mining Data works that pairing cell by cell.
Who builds on world mining data?
Ranked by how directly one row settles the day job:
- Metals and mining equity analysts see portfolio effects a single-metal view misses - sixty-five commodities in one consistent frame, with r/e/p flags telling them which years are safe to anchor a valuation on.
- Strategy and market-entry consultants turn "is this supply base concentrated?" into a number: market-share ranks, cumulative shares and HHI indices per commodity, deck-ready. See market researchers and analysts.
- Investors and quant researchers screen supplier concentration before it prints in earnings, pairing the production cube with price benchmarks from January 1960 onward - the workflow continues on investors and quants.
- Procurement and sourcing teams watch country-level production trends signal where supply of anything from cobalt to potash is concentrating, ahead of contract negotiations.
- Policy and trade economists cite a ministry-compiled tally spanning 168 countries precisely because its compiler sits outside every market it measures; the method is worked through in citation-grade research.
Use cases cluster into supply-concentration screening, market sizing and cross-country benchmarking - the first built on ranking rows joined to production history, the second on the cube rolled to income bands and blocs, the third on the long continent series.
Where to go next
Start with the record behind this article: the World Mining Data product page carries sample rows, the full field dictionary, coverage chips and delivery options, and a sample comes back shaped to the commodities and countries you name. Then move outward:
- Chile copper production statistics - the same ledger's biggest single-market story, drilled to operator level.
- Mineral reserves and resources data - what sits in the ground rather than what came out of it, deposit by deposit and filing by filing.
- Diversified metals & mining data hub - all 16 primary records with coverage, granularity and use cases side by side.
- Best diversified metals & mining datasets - the ranked shortlist and why the leaders lead.
- MSHA mine data vs World Mining Data - global production breadth against US mine-level safety records.
When you are ready to stop assembling and start analyzing, get a sample matched to your own commodity list and judge the columns, not the promise.
| field | type | definition | example |
|---|---|---|---|
| record_type | string | Which panel a row belongs to - production quantities by country, or the per-commodity world ranking with shares and concentration measures. | wmd_production_by_country |
| commodity_sheet | string | Mineral raw material the row reports, named for the commodity with its content basis appended wherever chemistry differs from elemental tonnes. | Copper · Chromium (Cr2O3) |
| country | string | Producer country; rollable up to continent, development status, income band, political-stability class or economic bloc. | Chile |
| unit | string | Unit of measure for that commodity - metr. t, kg, ct or Mio m3 - carried in-row beside every value instead of living in a header. | metr. t |
| y_<year> | number | Mine production of recoverable mineral content for the stated year; the country panel spans the five most recent editions-years. | 5505890 |
| quality_flag | string | Observation status riding in-band: r reported, e estimated, p provisional. Estimated and provisional values revise between editions. | r |
| rank_<year> | number | World rank of the producing country for that commodity-year, with the prior-year rank carried alongside in parentheses when unchanged. | 1 · (1) |
| production_<year> | number | Absolute tonnage behind a ranking row, stated on the same recoverable-content basis as the quantity panel. | 3100234 |
| share_pct | number | Country share of world production for the commodity-year; descending down the ranking, so a broken sort exposes a bad unpivot immediately. | 24.03 |
| hhi_contribution | number | Country's contribution to the commodity's Herfindahl-Hirschman concentration index; summing the column rebuilds the measure. | 577.51 |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
World Mining Data
World Bank Commodity Markets (Pink Sheet) Data
Cochilco Chilean Copper Statistics Data
Eurostat Mining and Quarrying Statistics
Natural Resources Canada - Mineral Statistics Data
Minerals Council South Africa - Facts & Figures
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What does the World Mining Data production panel contain?
Mine production of recoverable mineral content - not run-of-mine ore - for 65 mineral raw materials produced in 168 countries, organised into iron and ferro-alloy metals, non-ferrous metals, precious metals, industrial minerals and mineral fuels. Country-by-commodity detail covers the five most recent years (2020-2024 in the 2026 edition), and each observation carries a quality flag marking it reported, estimated or provisional.
How far back does world mining data go?
Two depths, and the difference matters. World totals by continent reach back to 1984; per-country, per-commodity detail covers only the five most recent years. A forty-year country panel does not exist in any single vintage, so longer histories are stitched across editions with revisions reconciled - scoped as an additional field set on the same feed when you request it.
Which dataset pairs with world mining data for prices?
The World Bank Commodity Markets (Pink Sheet) - monthly and annual benchmarks for 70+ commodities including aluminium, iron ore, copper, nickel and gold, running from January 1960 with 805 monthly observations per commodity, stated in nominal and constant-2010 dollars. Pairing the production cube with the price layer turns a volumes view into a values view, and both arrive through the same channel.