EPA Envirofacts Web Services (DMAP REST + GraphQL API)

Datadory delivers environmental-facilities-services data covering the EPA Envirofacts program tables behind the DMAP REST and GraphQL interfaces: TRI facility registers, SEMS Superfund sites with NPL status, ICIS compliance activity reports, plus RCRAInfo, GHG, NEI, NGGS, SDWIS and RADNET tables. Every field definition below was verified against live row output. Delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
United States - program tables carry state, county, ZIP and decimal-degree coordinate columns down to the individual facility or site
How far back
Varies by program table: ICIS and enforcement tables carry fiscal-year date fields, SEMS carries NPL and cleanup status dates, TRI reporting years run 1987 to present
How fine
Row-level program database records - facility, activity, chemical or monitoring row - never pre-aggregated statistics

What is the EPA Envirofacts Web Services dataset?

It is the queryable layer over the databases behind EPA's Envirofacts system, served through the agency's Data Management and Analytics Platform. Two interfaces reach the same tables: a DMAP-EF RESTful data service that takes URL-patterned requests shaped [table]/[column][operator][value]/[join]/[first]:[last]/[sort]/[format], and a GraphQL-style query service supporting field selection, aliases, filtering, joins, subqueries, pagination, ordering, grouping and variables.

Nine program families sit behind those interfaces. TRI covers toxic-release facilities; ICIS covers compliance activity reports; SEMS covers Superfund sites; RCRAInfo covers hazardous waste handlers; GHG covers greenhouse gas reporting; NEI covers the emissions inventory; NGGS, SDWIS and RADNET round out groundwater, drinking water and radiation monitoring. Tables are addressed as [program].[table] - tri.tri_facility, sems.envirofacts_site, icis.icis_activity_report - so one mental model spans all nine.

The operators matter as much as the tables: equals, notEquals, lessThan, greaterThan, beginsWith, endsWith, contains, like, in and notIn, combinable with /and/ and /or/, with case-insensitive text comparison and left joins between tables. In practice that means server-side filtering - you ask for Virginia ZIP codes or Region 05 source-test activities, not the whole nation. A single pull tops out at 100,000,000 records under a hard 15-minute cutoff, which bounds any one extract while leaving room for very large sweeps.

What do sample rows look like?

Each record family has its own shape, and the three below were captured live during research:

_table           : tri.tri_facility
tri_facility_id  : 24210TRTBN14378
facility_name    : TRI-TUBE INC
street_address   : 14378 ENTERPRISE RD
city / state     : ABINGDON / VA 24210
coordinates      : 36.740556 / 81.900556

_table           : sems.envirofacts_site
site_id          : 0406480
epa_id           : FLD984216630
name             : WASHAC INDUSTRIES
city / state     : ST. AUGUSTINE / FL
npl_status       : Not on the NPL
coordinates      : 29.957778 / -81.350556

_table           : icis.icis_activity_report
activity_name    : Region Code: 05 Action Type: 37 SOURCE TEST OBSERVED
status_desc      : (null)
begin_date_fy    : (null)

Read them together and the design shows: the TRI row is a manufacturing facility pinned to a street corner in southwest Virginia, the SEMS row is a Florida industrial site whose NPL status answers the first question any acquirer asks, and the ICIS row is a federal source-test observation tagged to Region 05. Same corpus, three regulatory lenses.

What fields does the dataset include?

Sixteen fields carry the analytical weight across the three sampled families, and every definition above was checked against live row output - the dictionary describes what actually arrives. Three structural facts are worth knowing before you model it.

First, identifiers stack rather than duplicate: the TRI facility ID keys release reporting, the FRS epa_registry_id links a site across EPA programs, and SEMS carries both an internal site_id and an epa_id. Joins between program families run on those identifiers, not on fuzzy name matching. Second, geography is uniform: city, state, ZIP and decimal-degree coordinates appear across the TRI and SEMS families, so radius queries and territory cuts work without per-program gymnastics. Third, nulls are information - the ICIS activity row above carries empty status and fiscal-year begin fields, which is normal on some activity types rather than a data fault.

How is the data covered across geography, time and granularity?

Geography - United States-wide, with state, county, ZIP and coordinate columns running through the program tables. The same facility can be located three ways, which matters when you are screening by drive-time rather than by political boundary.

Temporal - program-dependent by design. ICIS and enforcement tables carry fiscal-year date fields such as actual_begin_date_fy; SEMS carries NPL and cleanup status dates; the TRI tables hold reporting years from 1987 to the present. No single timeline covers all nine programs - which is precisely why delivery cadence is a subscriber decision here: daily if you watch enforcement activity break, weekly for portfolio reviews, hourly for event-driven work.

Granularity - one row per facility, per site, per activity, per chemical or monitoring observation. Nothing arrives pre-aggregated, so national totals, state rankings and site-level timelines all come out of the same rows. For scale: the companion ECHO Exporter alone spans more than 1.5 million regulated facilities, and this corpus reaches the program databases underneath that universe.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Who uses this data, and for what?

  • Site due diligence. NPL status and non-NPL cleanup language turn 'is there a Superfund problem here' into a query instead of a site visit. Coordinates make a radius screen around a target parcel a one-line filter.
  • ESG and supply-chain exposure. parent_co_name rolls facility-level releases up to corporate groups, so a holding company's environmental footprint aggregates without manual entity mapping. The FRS registry ID ties the same site across EPA programs.
  • Enforcement monitoring. ICIS activity reports - region-coded, action-typed, fiscal-year dated - give compliance teams a stream of who got inspected, tested and cited, filterable server-side to their own sectors and states.
  • Insurance and lending models. NPL proximity, cleanup status and coordinates feed liability scoring at address level rather than ZIP level.
  • Market sizing for remediation services. Counting sites by NPL status, state and cleanup type sizes the addressable market for state-lead and other non-NPL cleanup work directly from the status fields.
  • Journalism and academic replication. Row-level public-regulatory records mean any published total can be recomputed independently - the same rows EPA's own tools summarize.

Which personas get the most value?

Data scientists get row-level regulatory records clean enough to model on - see environmental & facilities services data for data scientists. Developers and builders get a documented query surface over nine program families without building nine integrations - see developers builders use cases. Journalists and academics get reproducible official numbers straight from program tables - see journalists academics use cases. Competitive intelligence teams track which facilities and corporate parents appear in new compliance activity - see competitive intel product teams use cases. Market researchers size remediation and compliance-service demand from status and geography fields - see market researchers use cases. Sales and growth teams build territory lists from program, state and ZIP filters - see sales growth teams use cases.

Why request a sample of this dataset?

Because the useful test is a join across program families, not a glance at one table. A sample comes back cut to the states, ZIPs, programs or corporate parents you name, with identifiers intact so you can verify immediately that a TRI facility resolves to its FRS registry record and that SEMS sites near your targets carry the NPL statuses you expect. Test whether your site list, customer accounts or portfolio holdings key cleanly onto EPA's identifiers before anything reaches production. Start small: get a sample of this dataset scoped to the slice you will actually use, or browse the best environmental facilities services datasets for companions.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary — sixteen verified fields across the TRI, SEMS and ICIS record families
FieldTypeDefinitionExample
tri_facility_idstringTRI facility identifier in the tri.tri_facility table - the stable key tying a plant to its release reporting history.24210TRTBN14378
facility_namestringRegistered facility name as reported in tri.tri_facility.TRI-TUBE INC
street_addressstringStreet address column in tri.tri_facility, sufficient for geocoding and parcel matching.14378 ENTERPRISE RD
city_namestringCity name present in both the TRI facility and SEMS site tables, so the two families share a geography spine.ABINGDON
state_abbrstringTwo-letter state abbreviation, filterable in tri.tri_facility.VA
zip_codestringZIP code column in tri.tri_facility - the fastest cut for territory-level screening.24210
pref_latitudenumberPreferred latitude in decimal degrees from tri.tri_facility.36.740556
pref_longitudenumberPreferred longitude in decimal degrees from tri.tri_facility; pairs with pref_latitude for radius queries.81.900556
epa_registry_idstringFacility Registry Service identifier in tri.tri_facility - the cross-program ID that links a site across EPA systems.110001134022
parent_co_namestringParent company name in tri.tri_facility, rolling facility exposure up to corporate groups.NA
activity_namestringCompliance activity description in icis.icis_activity_report, embedding region code and action type in the string itself.Region Code: 05 Action Type: 37 SOURCE TEST OBSERVED
actual_begin_date_fydateFiscal-year activity begin date in icis.icis_activity_report; null on some activity rows.
site_idstringSEMS internal site identifier in sems.envirofacts_site.0406480
epa_idstringEPA identifier for a Superfund site in sems.envirofacts_site.FLD984216630
npl_status_nameenumNational Priorities List status description in sems.envirofacts_site - the field that separates listed sites from the rest.Not on the NPL
non_npl_status_namestringCleanup status description for sites outside the NPL in sems.envirofacts_site, carrying state-lead and other activity designations.Other Cleanup Activity: State-Lead Cleanup

Coverage chips

DimensionCoverage
GeographyUnited States - state, county, ZIP and decimal-degree coordinate columns across the program tables
TemporalProgram-table dependent: fiscal-year date fields in ICIS/enforcement, NPL and cleanup status dates in SEMS, TRI reporting years 1987-present
GranularityOne row per facility, site, activity, chemical or monitoring observation - never aggregated
Record families3 sampled and verified - TRI facility registers, SEMS Superfund sites, ICIS activity reports - within 9 exposed program families

Additional fields available on request

Field groupNotes
Response envelope fieldsResult-count and status-message fields included in file deliveries on request.
Extended TRI and SEMS columnsRemaining columns of each program table beyond those exercised during verification, confirmed against your scope with the sample.
Chemical and monitoring attributesChemical, release-medium and monitoring-detail attributes from the program tables, delivered with the sample where they resolve.

Questions buyers ask

Which EPA programs are inside the Envirofacts dataset?

Nine programs are exposed: TRI (Toxics Release Inventory), ICIS compliance monitoring, SEMS Superfund sites, RCRAInfo hazardous waste handlers, GHG greenhouse gas reporting, NEI emissions inventory, NGGS, SDWIS drinking water and RADNET radiation monitoring. Each is addressed as its own schema.table pair, so a single delivery can hold facility registers, activity reports and site status records side by side.

What fields does the Envirofacts dataset include?

Verified fields span three record families. TRI facility rows carry the TRI ID, facility name, street address, city, state, ZIP, preferred latitude and longitude, the FRS registry identifier and parent company name. SEMS site rows carry internal site IDs, EPA IDs, NPL status descriptions and non-NPL cleanup status language. ICIS activity rows carry activity names, status descriptions and fiscal-year begin dates.

How granular is the data?

Row-level throughout: one record per facility, per site, per compliance activity, per chemical or monitoring observation - never pre-aggregated statistics. That is what separates this corpus from country-level waste indicators, and it means you can count, filter and join at whatever rollup your analysis needs.

How far back does the data go?

It depends on the program table. ICIS and enforcement tables carry fiscal-year date fields such as activity begin dates; SEMS carries NPL listing and cleanup status dates reaching back decades; the TRI tables hold reporting years from 1987 onward. Datadory deliveries accumulate across pulls, so your warehouse builds history even where a single snapshot would not.

Can I get TRI, SEMS and ICIS joined into one facility view?

Yes - that is the standard request. Records cross-reference through identifiers like the TRI facility ID, the FRS registry identifier and the SEMS site/EPA IDs, so one facility can pick up its release reporting, cleanup status and compliance activities in a single keyed row set. Ask for the join in your sample request and we return it pre-built.

Who uses Envirofacts data?

Environmental consulting firms screening sites for acquisition due diligence, ESG analysts tracing parent-company exposure to reported releases, insurers pricing environmental liability by NPL proximity, journalists reproducing official enforcement numbers, and sales teams targeting facilities by program, state and ZIP.

Notes on this record

  • Definitions verified against live rows All sixteen dictionary fields were checked against actual row output during research - a Virginia TRI facility, a Florida SEMS site and a Region 05 ICIS activity report. The page describes arriving data, not brochure data.
  • Only GraphQL-style route into EPA program tables Of the 1,744 datasets Datadory catalogs, this is the only EPA record exposing GraphQL-style querying over program tables - field selection, joins, subqueries and pagination in one query shape alongside the REST pattern.
  • Identifiers stack across programs TRI facility IDs, the FRS registry identifier, SEMS site IDs and EPA IDs all appear in the same corpus, so one facility picks up release reporting, cleanup status and compliance activity through keyed joins rather than name matching.
  • Table naming does not follow acronyms Tables are addressed as [program].[table], and the spelling does not always match the program acronym - one guessed handler-table name was rejected as not found during research. Datadory resolves exact table names before any delivery ships, so your scope never dies on a typo'd table.
  • Pair it with bulk companions For scheduled everything-at-once ingestion, EPA ECHO Data Downloads covers more than 1.5 million regulated facilities in bulk form. This corpus complements it: filtered, row-level program records on your cadence instead of bulk snapshots on a fixed schedule.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing