U.S. Energy Information Administration

EIA Petrochemical Feedstock & Refinery Olefin Production Data

Datadory delivers EIA Petrochemical Feedstock & Refinery Olefin Production data covering every barrel of naphtha US refineries net out for petrochemical feedstock use - the input line behind American ethylene economics. Sixteen series per layout run the national total back to 1983 and PADD detail from 1993, monthly and annually.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
United States total plus PADDs 1-5 - East Coast, Midwest, Gulf Coast, Rocky Mountain, West Coast - and refining districts within them, from Texas Inland to North Louisiana-Arkansas
How far back
United States series runs 1983-2025 annually; regional series generally start 1993, some shorter, at least one (Appalachian No. 1) discontinued
How fine
monthly and annual national/regional totals per product; no facility, company or plant level anywhere in the family

What is the EIA Petrochemical Feedstock & Refinery Olefin Production Data?

It is the federal measurement of one specific flow: naphtha that US refineries and blenders produce net for sale as petrochemical feedstock, as opposed to naphtha that stays in the fuels pool. That split matters because feedstock naphtha is the principal light-liquid input to American steam cracking, and its availability is a first-order driver of ethylene and bulk olefin production economics.

The U.S. Energy Information Administration publishes the family under its petroleum refining statistics, compiled from the Petroleum Supply Monthly and the agency's weekly and monthly refinery surveys. One product, many cuts of geography: a United States total (series key MNFRPUS2), five Petroleum Administration for Defense Districts, and the refining districts nested inside PADD 2 and PADD 3 - Indiana-Illinois-Kentucky, Minnesota-Wisconsin-Dakotas and Oklahoma-Kansas-Missouri in the Midwest; Texas Inland, Texas Gulf Coast, Louisiana Gulf Coast and North Louisiana-Arkansas on the Gulf. A discontinued Appalachian No. 1 district rounds out the set. Each series carries its own MNFRP-prefixed key, which is what makes district-level joins mechanical rather than manual.

What do the rows actually look like?

One row per series per period, values in thousand barrels per day or thousand barrels depending on the view. The national annual series reads 126 thousand barrels per day for 2022, 131 for 2023, 123 for 2024 and 142 for 2025 - a nineteen-unit jump between 2024 and 2025, roughly fifteen percent more naphtha leaving refinery gates for crackers in a single year. That kind of move is exactly what feedstock-demand models exist to catch.

Regional rows arrive identically shaped. Midwest (PADD 2) printed 13 thousand barrels per day in 2024 against the national 123 - the Gulf Coast carries most of the feedstock weight, and the district columns make that concentration visible rather than anecdotal. Cells a district did not report arrive legend-flagged with the publisher's -, --, NA and W markers, never silently zero-filled, so summing across districts cannot manufacture volume that does not exist.

What fields does each record carry?

Sourcekey is the spine - the MNFRP-family identifier that distinguishes the US total from each PADD and district series, and the natural join key against any procurement, trading or model table you bolt on. Date carries the reporting period, end-of-period dated on annual rows. The production value sits in the unit the view selects: thousand barrels for cumulative volume, thousand barrels per day for rate.

Two dimensional fields complete the row. Product is constant here - naphtha for petrochemical feedstock use, product code EPPPN - and exists so this family can sit beside sibling product series without ambiguity. Area records the layout orientation the publisher offered, product-by-area or area-by-product, which sounds cosmetic until a reshaped pivot quietly drops which cut a figure came from. Delivered rows normalize region into its own field so PADD-level filtering is a WHERE clause, not spreadsheet surgery.

Where does the coverage sit?

Geographically: the United States total, then PADDs 1 through 5 - East Coast, Midwest, Gulf Coast, Rocky Mountain, West Coast - then the districts inside them, with the Gulf Coast split finest because that is where the crackers concentrate. Temporally, the national series reaches from 1983 through 2025, forty-three annual observations, with monthly views layered underneath. Regional series generally begin in 1993; some run shorter, and Appalachian No. 1 is discontinued, which is why every delivered table carries its series status rather than padding gaps with zeros.

Granularity is deliberately bounded: monthly and annual totals per product per geography. There is no facility level, no company attribution, no plant throughput anywhere in the family - questions at that altitude belong to companion sources such as the EIA Refinery Capacity Report, which pairs cleanly against these volumes.

How is the data delivered?

Through Datadory, on your terms: API, files, or landed directly in your warehouse. Daily, weekly, or hourly - the cadence is your call, not the publisher's. Rows arrive normalized with region, product, unit and frequency as explicit fields, legend-flagged where the source withholds, and documented against the field dictionary above so the first join works on day one.

Who uses this data, and for what?

Feedstock and margin analysts read the national and Gulf Coast series as the supply side of US ethylene economics - when net feedstock production moves fifteen percent in a year, cracker utilization and derivative margins move with it. Energy market researchers track the PADD splits to see where naphtha demand competes with gasoline blending inside the refinery slate.

Corporate development teams use the district detail when siting or valuing Gulf Coast and Midwest petrochemical assets, pairing feedstock availability against the capacity report. Procurement and supply-chain planners at naphtha consumers benchmark their own purchases against measured national and regional flows. Journalists covering the American petrochemical build-out cite the annual prints because they are federal measurements, not trade estimates. Quant desks fold the series into chemical-equity factor models as a hard supply signal.

Which personas get the most value?

Investors and quants get a federal, monthly-resolution supply input for chemicals coverage. Competitive intelligence and product teams benchmark whether feedstock conditions in a district support the expansion story a competitor is telling. Market researchers and consultants anchor petrochemical demand studies in measured flows rather than consultant estimates. Data scientists and ML engineers inherit a clean panel - series key, period, geography, value - ready for feature pipelines. Journalists and academics get citable federal numbers with four decades of history.

Why request this through Datadory?

Because the raw family is a lattice of parallel series views - product-by-area, area-by-product, four unit-and-frequency combinations per layout - and reconstructing one honest table from it is easy to get subtly wrong. Datadory hands you the normalized version: one row per series per period, region and unit explicit, withheld cells flagged, district status documented. Name the PADDs, districts and periods you need and a sample comes back shaped like your model expects, with companion sources - the wider petroleum portal, the refinery capacity report - joinable on the same geographic keys.

Which notes pair with this dataset?

The cards below sit closest to this dataset in our catalog: the broader petroleum portal that frames these series, the capacity report that supplies the denominator, property references for the molecules themselves, and the vertical ranking that shows where this family lands among commodity chemicals sources.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - the six working fields on every delivered row; examples taken from the 2025 United States annual print (Sourcekey MNFRPUS2)
FieldTypeDefinitionExample
SourcekeystringEIA series identifier - the join-friendly key, MNFRP-prefixed, that separates the United States total (MNFRPUS2) from each PADD and refining-district series.MNFRPUS2
DatedateReporting period for the row. Annual rows carry end-of-period dates; monthly rows carry the month, so a year of monthly observations stacks twelve rows where the annual view carries one.2025-12-31
Refinery and Blender Net ProductionnumberThe value itself: net production of naphtha for petrochemical feedstock use for the row's series and period, in the chosen unit - thousand barrels or thousand barrels per day.142
RegionstringGeography the row reports: the United States total, one of PADDs 1-5 (East Coast, Midwest, Gulf Coast, Rocky Mountain, West Coast), or a refining district such as Texas Inland or Louisiana Gulf Coast. Raw layouts carry one column per region; delivered rows normalize it into this field.Gulf Coast (PADD 3)
ProductstringProduct dimension, constant across this family: naphtha for petrochemical feedstock use, carried internally as product code EPPPN.Naphtha For Petrochemical Feedstock Use
AreastringLayout dimension - whether the publisher's view pivots product-by-area or area-by-product. Preserved on delivery so a reshaped pivot never loses which cut a number came from.U.S. By Product

Questions buyers ask

What is the EIA Petrochemical Feedstock & Refinery Olefin Production Data dataset?

A family of US government time series measuring refinery and blender net production of naphtha for petrochemical feedstock use - naphtha sold onward as cracker feed rather than blended into fuels. It covers the United States total, PADDs 1-5 and the refining districts within them, in monthly and annual views, in thousand barrels and thousand barrels per day, with the national series running 1983-2025.

What is naphtha for petrochemical feedstock use?

Light naphtha destined for steam cracking into ethylene, propylene and other olefins rather than for gasoline blending. Refineries report the split, which makes this series a direct read on how much cracking feedstock the US refining system is supplying. Because feedstock cost and availability dominate ethylene economics, the series functions as a leading indicator for bulk petrochemical margins.

What are PADDs and why do they matter here?

Petroleum Administration for Defense Districts - five legacy supply regions (East Coast, Midwest, Gulf Coast, Rocky Mountain, West Coast) the energy statistics system still uses as its standard geographic cut. They matter because petrochemical capacity concentrates in PADD 3, the Gulf Coast: the district-level detail inside PADD 3 - Texas Inland, Texas Gulf Coast, Louisiana Gulf Coast, North Louisiana-Arkansas - is where feedstock availability meets cracker locations.

How far back does the data go?

The United States annual series spans 1983 through 2025 - forty-three years of history, with monthly views layered beneath it. Regional series generally begin in 1993, though some districts run shorter histories and at least one, Appalachian No. 1, is discontinued. Every delivered table documents each series' start, end and active status so gaps are known quantities rather than surprises.

Does the dataset cover ethane, propane or other feedstocks?

This cataloged family covers naphtha for petrochemical feedstock use. The publisher maintains parallel product families for other hydrocarbon streams in the same refining statistics framework, and those can be folded into a sample on request - name the products and PADDs you need. What the family never includes is facility-level or company-level detail; geography stops at the refining-district line.

Is there plant-level or facility-level detail?

No. The series measure net production aggregated to national, PADD and refining-district level - one product, many geographies, no individual refineries named and no companies attributed. Facility questions pair better with the EIA Refinery Capacity Report, which lists individual refineries and their capacities, joinable against these volumes by PADD and district.

What do the legend symbols in the data mean?

Four markers appear where a value is absent: single dash, double dash, NA and W, indicating withheld or not-reported figures. They mean a district did not report a usable value for that period - not zero production. Delivered tables preserve the flags rather than filling gaps, which keeps sums and averages across districts honest and prevents phantom volume entering a model.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing