OSMRE Mining Regulation & Reclamation Data
Datadory delivers OSMRE mining regulation reclamation data covering US surface coal mining under SMCRA: the e-AMLIS abandoned-mine inventory behind a published 63-field dictionary, Priority 1/2 and 3 problem taxonomies, $14.233 billion collected and $6.569 billion granted from the AML fund, AVS violator records, and 275,000-plus mine maps.
- geo
- United States - 24 primacy states plus federal-program states and Indian tribes with legacy coal mining under SMCRA Titles IV and V
- How far back
- Abandoned-mine impacts tracked continuously since 1977 with fee authority through September 30, 2034; annual reports FY2017-FY2020; grant distributions FY2023-FY2026; mine maps back to the 1790s
- How fine
- Problem-area and problem-component level keyed on AMLIS KEY within planning units and watersheds; program aggregates in the fund-status, grant-distribution and annual-report layers
What is the OSMRE Mining Regulation & Reclamation Data dataset?
It is the United States' official ledger of coal mining's regulatory record and reclamation debt, kept by the Office of Surface Mining Reclamation and Enforcement - the Interior Department agency that administers the Surface Mining Control and Reclamation Act (SMCRA) of 1977.
Two programs give the record its spine. Under Title IV, the e-AMLIS Abandoned Mine Land Inventory System catalogs land and water impacted by pre-1977 coal mining: every problem area carries location, problem type, SMCRA priority class, extent, census-derived population risk, stream miles, impounded acres and construction cost, split across unfunded, funded and completed statuses. Under Title V, the record turns to active-mine regulation - the Applicant Violator System of permittee, operator and unabated-violation records used for permit eligibility under SMCRA 510(c), plus hydrologic assessments and bonding guidance.
The money side is documented just as hard. The AML Reclamation Fund has collected $14.233 billion since fees began on August 3, 1977 and distributed $6.569 billion in fee-based grants; annual grant distributions run FY2023-FY2026; the AMLER program has drawn more than $1 billion in appropriations since FY2016 with state-by-state allocations; and an $11.293 billion funding extension runs through 2034. Get a sample of this dataset and the slice you name comes back as typed rows, not document excerpts.
What do the sample rows look like?
One e-AMLIS problem area arrives shaped like this - the values are the agency's own worked dictionary examples, printed so the shape is visible before live records arrive:
# one record = one problem area within a planning unit
dataset : osmre_eamlis_problem_area
amlis_key : AL000001
pa_name : New Lexington
state : AL
county : Tuscaloosa
quadrangle : New Lexington
watershed : Upper Black Warrior
huc_code : 3160112
mining_type : S # surface / underground / both
priority : 3 # 1-2 safety threats, 3 environmental
problem_type : SA - P3 Spoil Area
census_population : 3581
latitude : 33.56361111
longitude : -87.651388889
# reclamation-project leg on the same key
project_name : Deans Ferry Acid Mine Drainage Remediation
stream_miles : 2
impounded_acres : 1
census_risk : 622Live records swap those placeholders for real sites - named highwalls, burning refuse banks, stream miles reborn - while keeping the column shape. Every geographic attribute derives from the recorded coordinates, so one row supports map pins, watershed roll-ups and county tabulations without a second lookup table.
Which fields does the dictionary define?
A published 63-field data dictionary covers e-AMLIS - data types, sizes, descriptions and a worked example for every column - which is why this catalog scores the record's field definitions as verified. The table below carries the fields that structure most builds.
Two design choices matter before you model it. First, geography is deliberately redundant: county, quadrangle, watershed, FIPS code and HUC code all derive from the problem-area coordinates, so the same record joins cleanly at whatever geography your analysis speaks. Second, the extent-and-cost block repeats across the funding lifecycle - unfunded, funded and complete - turning 'what reclamation is still owed' into a single filter rather than a reconciliation project. Planning-unit names, congressional district, ore type, the five-date lifecycle set and the land-ownership percentages round out the dictionary; anything beyond the core columns below joins the feed as additional fields on request.
Where does coverage run, and at what grain?
- Geo: United States - 24 primacy states plus federal-program states and Indian tribes carrying legacy coal mining under SMCRA Titles IV and V.
- Temporal: the abandoned-mine inventory tracks impacts from pre-1977 mining continuously since SMCRA's 1977 enactment, with fee authority extended to September 30, 2034; annual reports are posted for FY2017-FY2020 and grant distributions for FY2023-FY2026; the National Mine Map Repository holds maps reaching back to the 1790s.
- Granularity: problem-area and problem-component level keyed on AMLIS KEY, nested inside planning units and watersheds; program-level aggregates sit in the fund-status, grant-distribution and annual-report layers.
One structural note: the inventory layer and the program-aggregate layer answer different questions - exposure versus dollars moved - and naming which one your build needs returns the right sample shape first.
How is the data delivered through Datadory?
API, files, or your warehouse. Daily, weekly, or hourly.
Pick the channel your stack already speaks and set the cadence to match the decision being fed. Site-inventory rows slot into warehouse loads beside your operational data; the fund-status and grant layers reward a pull whenever a budget cycle or appropriation moves; the violator records earn their keep as discrete events worth catching the day they change. Cadence changes are a settings conversation, not a re-integration project.
Fields beyond the core columns - deeper geographic cuts, the full ownership breakdown, lifecycle dates - ride the same channel once scoped. Name them with your sample and they join the feed instead of spawning a second pipeline.
Who builds on this dataset?
Ranked by how directly one row settles the day job:
- Remediation engineers and reclamation contractors. The unfunded-extent and unfunded-cost columns are a national pipeline of work still owed - problem types, acreage and construction estimates, before anyone posts a bid.
- Policy researchers and economists. $14.233 billion collected against $6.569 billion granted, AMLER allocations by state since FY2016, and an $11.293 billion authorization through 2034 - a citable accounting of who pays and who benefits.
- Investors and quants. Legacy liability has coordinates: population-exposure counts, stream miles and priority classes attach directly to the landscapes around mining assets.
- Journalists and academics. Highwalls, mine fires and portals are defined problem types with communities counted beside them - a regulatory record that quantifies its own hazards.
Which personas get the most value?
Market researchers and consultants get the remediation economy quantified - problem taxonomies, costs and fund flows in one addressable place. Investors and quants get location-grade legacy liability to overlay on asset screens. Journalists, academics and students get the original measurement - hazard classes and exposed populations straight from the regulator's own dictionary rather than a redrawn summary. Data scientists get a coordinate-keyed panel whose 63-column dictionary was published in full, so feature engineering starts at analysis instead of schema archaeology.
What should you know before requesting a sample?
Three things, stated up front. First, this is the regulatory and reclamation record, not production tonnage - if your model needs output volumes, pair it with MSHA mine reports rather than stretching these columns. Second, the inventory is site-first: records key on individual problem areas, so national totals are yours to aggregate while the program-dollar layer arrives as report-style aggregates. Third, the dictionary is complete and public-facing - 63 columns with definitions and examples - so scope conversations happen against real documentation, not guesses.
Get a sample of this dataset, or browse the diversified metals mining data hub for the rest of the industry.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
amlis_key | string | State or Tribe abbreviation followed by six auto-generated or manually assigned integers identifying the problem area. | AL000001 |
pa_name | string | Name given to the problem area by a State or Tribe. | New Lexington |
state | string | Two-letter abbreviation for the State or Tribal Nation of the record. | AL |
county | string | County auto-generated from the latitude/longitude entered for the problem area. | Tuscaloosa |
quadrangle | string | USGS topographic quadrangle name derived from the problem-area coordinates. | New Lexington |
watershed | string | Watershed Boundary Dataset name derived from the problem-area coordinates. | Upper Black Warrior |
huc_code | string | Hydrologic unit code identifying the drainage area containing the problem. | 3160112 |
priority | enum | SMCRA 403(a)(2)/411(c)(2) priority class 1-5; Priorities 1 and 2 mark health-and-safety threats, 3 marks environmental restoration. | 3 |
problem_type | enum | Abbreviation and name of the specific on-the-ground feature - 18 Priority 1/2 types, 14 Priority 3 types. | SA - P3 Spoil Area |
mining_type | enum | Type of pre-SMCRA mining in the problem area: surface, underground or both. | S |
ore_type | string | Ore or commodity type for non-coal reclamation sites, e.g. Bentonite, Cinnabar, Clay, Copper. | C |
census_population | integer | Population derived from Census Bureau block-group counts around the problem area location. | 3581 |
latitude | number | Angular distance north or south of the equator, carried as decimal(16,12). | 33.56361111 |
longitude | number | Angular distance east or west of the Greenwich meridian, carried as decimal(16,12). | -87.651388889 |
unfunded_cost | number | Estimated direct construction cost still needed on the unfunded share of the problem, carried with extent units and metric conversions. | |
funded_cost | number | Actual obligated direct construction cost on the funded share, with extent units and conversions. | |
completed_cost | number | Final direct construction cost on completed reclamation, with completed extent units. | |
project_name | string | Name assigned by a State, Tribe or OSMRE Field Office for an approved reclamation project. | Deans Ferry Acid Mine Drainage Remediation |
stream_miles | number | Total stream miles reclaimed on a water-related reclamation project. | 2 |
impounded_acres | number | Total impounded acres reclaimed on the reclamation project. | 1 |
census_risk | integer | Census-derived population exposure figure around the problem area, alongside state-supplied alternate estimates. | 622 |
Questions buyers ask
What exactly does the OSMRE dataset contain?
US surface coal mining regulation and reclamation under SMCRA: the e-AMLIS inventory of land and water impacted by pre-1977 mining with a published 63-field dictionary, AML fund status and grant distributions, Applicant Violator System permit records, hydrologic and bonding guidance, and the National Mine Map Repository of 275,000-plus mines.
What is the difference between Priority 1, 2 and 3 problems?
Priorities 1 and 2 are health-and-safety threats - 18 defined problem types such as dangerous highwalls, portals and underground mine fires. Priority 3 covers environmental restoration across 14 types, and historically flagged Priorities 4-5 round out the taxonomy. The priority class is a first-class column, so safety exposure filters in one predicate.
How well is the e-AMLIS schema documented?
Exceptionally: a published 63-field dictionary lists every column's data type, size, description and worked example, plus pages defining each Priority 1/2 and Priority 3 problem type. Modeling starts from real documentation rather than reverse-engineered headers.
What does the Applicant Violator System add?
Permittee, operator and unabated-violation records used to determine permit eligibility under SMCRA 510(c). It is the compliance counterpart to the abandoned-mine inventory - who holds permits, who has operated, and whose violations remain unabated - which is why diligence teams read it beside the reclamation ledger.
How much money has moved through the AML fund?
$14.233 billion collected since fees began on August 3, 1977, against $6.569 billion distributed in fee-based grants. Grant distributions run FY2023-FY2026, AMLER has drawn over $1 billion since FY2016 with state-by-state allocations, and an $11.293 billion authorization extends funding through 2034.
Can I evaluate real records before committing?
Yes - that is what the sample is for. Name the states, priority classes, problem types or program layer you care about, and real rows come back cut to that specification together with the full field dictionary, delivered via API, files, or your warehouse on a daily, weekly, or hourly cadence.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.