Datadory notebook

US refinery-level capacity data: every plant named, delivered as rows

Yes - Datadory delivers US refinery-level capacity data covering all 130 Form EIA-820 respondent sites as typed rows: atmospheric crude distillation capacity in barrels per calendar day and per stream day, downstream charge capacity for vacuum distillation, coking, catalytic cracking, hydrocracking, reforming and desulfurization units, and production capacity by product for every operating and idle plant, each named with operator, site, state and PADD district, in an annual census reaching back to 1994 - delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What is refinery-level capacity data?

The query sounds like a permissions question; the answer is a map of American refining. Refinery-level capacity data names the plant first - which company, at which site, in which state and PAD district - and attaches numbers to it: how many barrels per calendar day of crude that specific atmospheric distillation column can handle, how many per stream day when utilization is honest, and how deep the plant runs behind it. In the United States this record exists at full grain, and it is the EIA Refinery Capacity Report, the published face of the annual Form EIA-820 census.

The scope is unusual among energy surveys. The 2026 edition covers 130 respondents: every operating and idle refinery, new refineries still under construction and, since the 2024 survey, non-refinery operators of distillation, reforming, cracking, coking and hydrotreating units - across the 50 states, DC, Puerto Rico, the US Virgin Islands, Guam and other US possessions. Capacity figures are explicitly designated non-confidential at plant level, which is why a single company at a single site can appear beside a single number, while other parts of Form EIA-820 stay protected and never surface in the published tables.

Datadory delivers the whole record as typed rows keyed on the plant spine, scored 9 out of 10 on our quality rubric against a catalog-wide average of 7.81 across the 1,744 datasets we catalog; only 145 datasets earn better than 9. Get a sample cut to your districts and process units before anything else.

What does one row look like once delivered?

One row per refinery-product-capacity-measure after cleaning - illustrative values on the real skeleton:

CORPORATION              SITE          STATE  PADD  PRODUCT/UNIT        MEASURE                  QTY
------------------------------------------------------------------------------------------------------
Equistar Chemicals LP    Channelview   TX     3     ALKYLATES           Production capacity      22,370
Marathon Petroleum       Garyville     LA     3     CRUDE DISTILLATION  Capacity, bbl/cd        597,000
Valero Energy            Benicia       CA     5     COKING              Charge capacity, bbl/sd  52,000
Citgo Petroleum          Lake Charles  LA     3     REFORMING           Charge capacity, bbl/sd  68,000

Read what the columns do before reading what the numbers imply. The identity block - corporation, site, state, PADD - is the entire reason this record beats every aggregate feed: a ranking of Gulf Coast coking capacity is a group-by, not a reporting project. The two capacity bases are different quantities and must never be averaged together; calendar-day capacity spreads a year's running over 365 days, stream-day capacity assumes the plant runs flat out on the days it runs, and the spread between them at any large refinery routinely exceeds fifteen percent. And the third row's product column shows why the file flattens to ~3,336 rows: one refinery contributes dozens of rows across 34 process and product categories, so summing the QTY column without filtering MEASURE first produces a number that means nothing.

Which fields carry the weight?

The field dictionary is short and load-bearing, which is exactly why it gets abused:

  • Identity fields name the operating corporation, the physical site, the state and the PAD district. Site names are how the industry actually talks - Channelview, Baytown, Garyville - and they are the join key for everything else you know about a plant.
  • Atmospheric crude oil distillation capacity arrives twice, in barrels per calendar day and barrels per stream day. Pick the basis your denominator expects before computing anything.
  • Downstream charge capacity covers vacuum distillation, coking, catalytic cracking, hydrocracking, reforming and desulfurization units in barrels per stream day, with fresh feed input recorded alongside - the difference between the two is where conversion bottlenecks hide.
  • Production capacity by product spans the 34-category product list, alkylates through residual fuel, one row per refinery-product.
  • Status and change tables carry new, reactivated and permanently shutdown refineries, with shutdown history reaching back to 1990 - the consolidation story lives here rather than in the capacity totals.

The one trap worth naming out loud: the SUPPLY-style measure labels separate calendar-day crude capacity from stream-day charge capacity from next-year production capacity, and any ranking computed without filtering on the label mixes incompatible quantities. Datadory ships the measure split into named, typed columns so the filter is structural rather than remembered mid-analysis.

How far back does coverage reach, and where does it stop?

Temporal - an annual snapshot as of January 1 of each survey year, with editions running from 1994 through 2026: roughly three decades of January 1 readings for trend work on consolidation, capacity creep and regional concentration. Two years simply do not exist - the survey was not conducted for January 1, 1996 or January 1, 1998 - so any long-run panel carries two holes that have to be interpolated around or explicitly excluded. Auxiliary histories soften the edges: permanent-shutdown tables run from 1990, and a working and shell storage capacity series spans 1982 through 2010, useful when a capacity figure older than the first standalone edition is needed. Each new edition lands mid-year, measuring the previous January 1.

Geography - complete United States coverage organized by PAD district and state, with the five PADDs carrying the analytic weight: PADD 3 alone concentrates roughly half of national capacity, and district rollups ship as group-by operations because the district field rides on every plant row.

Granularity - plant-by-process-unit-by-product, the finest structural grain refining publishes anywhere. What it deliberately does not include is time within the year: no monthly runs, no quarterly utilization, no intramonth anything. Capacity is a stock; flows live elsewhere, and pretending otherwise produces utilization estimates that cannot survive review.

How does capacity differ from throughput and utilization data?

Forward views come from the third record in the stack: the EIA Short-Term Energy Outlook projects refinery balances and prices monthly through the end of the following calendar year, anchored on those same January 1 snapshots. Note the asymmetry with Canada, where the regulator publishes filed monthly throughput and capacity together for regulated pipeline systems - but those are pipelines, not refineries, and no equivalent continuous refinery-level series exists on either side of the border.

Where else does plant-level refining capacity exist outside the United States?

India is the closest analogue. PPAC's Monthly Ready Reckoner, from the Petroleum Planning & Analysis Cell, covers installed refinery capacity and crude processing by company, alongside product-wise production and consumption and gross refining margins, in a companion workbook of 107 worksheets with fiscal-year series from roughly FY2012-13 - but Indian plants only, and by company rather than by named site.

Beyond that, named-plant refining capacity thins out fast. Regional portals carry balances rather than registries, and European transparency platforms skew toward storage and prices: GIE AGSI+ reports daily filling levels for gas storage sites across EU member states plus the UK and Ukraine, while Eurostat's harmonized stock tables cover monthly levels for dozens of geographies. If the question is genuinely global plant-level refining capacity, the US census remains the deepest single public source, and everything else approximates toward it.

The comparison table below lines up the four records that matter for refining work; the head-to-head between the two EIA records lives in our Petroleum & Other Liquids Data Portal vs Refinery Capacity Report comparison.

Who builds on refinery-level capacity data?

Ranked by how directly the plant grain answers the day job:

  1. Sales & growth teams build ranked target lists of US refineries with capacities, process units and operating status for downstream outreach; the workflow ranks the slice in sales & growth teams.
  2. Market researchers & consultants benchmark refining capacity by PADD and state off a citable annual census reaching back three decades; see market researchers.
  3. Investors & quants compute utilization and capacity-creep assumptions for refining models, joining the annual snapshot to flow records before believing any management-deck trend line; see investors & quants.
  4. Data scientists & ML engineers train forecasting features on plant-level structure - capacity, unit mix, vintage of shutdowns - that aggregate feeds flatten away; see data scientists.
  5. Journalists, academics & students answer how-many-refineries and how-much-capacity questions from the official plant-level survey whose methodology is published end to end; see journalists, academics & students.

Why get refinery capacity data through Datadory?

Because the hard part was never obtaining the file - it is the tenth analysis built on top of it. Measure labels that mix two capacity bases in one column. Two missing survey years hiding inside a thirty-year panel. Site names spelled differently between vintages. A June release that rewrites last January's snapshot while your model still holds the old one. Each is survivable once; none is fun to re-solve in every notebook.

Datadory normalizes before delivery: the capacity measures split into named, typed columns so calendar-day and stream-day never average together, report-year vintage labeled on every observation so comparisons fail loudly instead of silently, entity resolution keeping site identities stable across editions, and the flow records pre-keyed to the same plant spine so utilization is a join rather than a reconciliation project. When the next annual edition lands, it arrives as new rows under the same dictionary - no re-integration.

Name the districts, process units and years when you request a sample and it arrives cut to that scope; the standing feed follows the same shape, so anything prototyped on the sample survives delivery intact.

Where to go next

Start with the EIA Refinery Capacity Report dataset page for the full field dictionary and to request a sample cut to your districts and process units. For the head-to-head that splits snapshot capacity from continuous flows, read EIA Petroleum & Other Liquids Data Portal vs EIA Refinery Capacity Report. Adjacent deep-dives: country level oil production statistics and global oil pipeline database GeoJSON. The Oil & Gas Refining & Marketing Data Guide scores all twenty records in the slice, the oil-gas-refining-marketing data hub holds the pooled scorecard, and the definition underneath it lives in the glossary entry for refinery capacity by process unit.

Coverage chips - geography, time and granularity for the EIA Refinery Capacity Report
DimensionCoverage
GeographyAll 50 states plus DC, Puerto Rico, the US Virgin Islands, Guam and other US possessions, organized by PAD district (PADD 1 through 5) and state
TemporalAs of January 1 of each survey year; editions 1994 through 2026 with no surveys for 1996 or 1998; shutdown tables from 1990; working and shell storage history 1982-2010
GranularityOne row per refinery-product-capacity-measure: corporation, site, state and PADD named; crude distillation in barrels per calendar day and per stream day; six downstream unit families plus product-level production capacity
MethodologyMandatory annual EIA-820 census covering operating and idle refineries, refineries under construction and, since 2024, non-refinery operators of major conversion units; capacity figures designated non-confidential at plant level

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Oil & Gas Refining & Marketing United States - all 50 states

EIA Refinery Capacity Report

Period · Supply

Oil & Gas Refining & Marketing United States - national

EIA Petroleum & Other Liquids Data Portal

duoarea · product · process …+4 more

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Is refinery-level capacity data publicly available in the United States?

Yes. The EIA Refinery Capacity Report - the published face of the annual Form EIA-820 survey taken as of January 1 - names atmospheric crude distillation, downstream charge and production capacities for about 130 respondent sites, with editions reaching back to 1994. Datadory delivers the same record as typed rows, one per refinery-product-capacity measure, alongside the flow records utilization work needs.

What does one refinery capacity row contain?

An operating company, a physical site, a state, a PAD district, a process unit or product, a capacity measure and a quantity - roughly 3,336 such rows across 34 process and product categories in the current edition. A sample row reads: Equistar Chemicals LP, Channelview site, Texas, PADD 3, ALKYLATES, production capacity, 22,370 barrels per stream day.

Does the report show actual crude runs and utilization by refinery?

No. Capacity is a stock measured once a year on January 1; actual runs are not part of it. Utilization work pairs the annual snapshot against flow records from the EIA Petroleum & Other Liquids Data Portal - roughly 53,300 series including weekly refinery inputs at national, PADD, state and individual-refinery granularity - and Datadory delivers both into one warehouse so the division happens on joined rows.

Which years are missing from the historical refinery capacity series?

The survey was not conducted for January 1, 1996 or January 1, 1998, so editions start in 1994 with two gaps; shutdown tables extend back to 1990 and working/shell storage capacity history spans 1982 through 2010. Every row Datadory ships carries an explicit report-year vintage label, so vintage mixing fails loudly instead of silently.

How is refinery capacity data delivered?

API, files, or straight into your warehouse - daily, weekly, or hourly. Capacity measures arrive split into named typed columns, report-year vintage labeled on every observation, and site identities resolved consistently across editions so rankings and panel builds need no cleanup pass. The field dictionary is unchanged between the sample and the production feed, so validation takes minutes.