Datadory notebook
EIA bulk download petroleum archives: the federal fuels catalog, delivered as rows
Datadory delivers EIA petroleum archive data covering every bulk series the U.S. Energy Information Administration publishes: weekly lower-48, PADD and state stocks observed back to April 2004, petroleum imports tracked to national, PADD, state, city, port and refinery level by crude grade and country of origin, and the natural gas production-through-stocks package - frozen in named versions like 81.0.0, thirty files near 1.6 GB, every resource checksummed in the manifest. Typed rows delivered daily, weekly, or hourly - your call.
1,744 datasets. Pick your catch.
Why build on a numbered version instead of a live feed?
Reproducibility is the whole argument. Each version is a frozen capture that ends at its assembly date, addressable by number, so a feature built against 81.0.0 returns identical figures months later - something a live feed cannot promise once the agency revises a week. The manifest ships a hash per resource, which reduces integrity checking to a one-line diff, and a pinned rebuild reproduces bit-for-bit.
Versioning also moves fast enough to matter: version 54.0.0, published January 25, 2026, held 26 files totalling about 1.46 GB; version 81.0.0, seven months later, holds 30 files near 1.6 GB. Chasing that by hand is a part-time job. Delivered through Datadory, the version pin travels with every extract - a dashboard, a model's training table and a board deck can cite the same capture - while newer captures slot in as their own numbered versions, never as edits underneath you. Currency becomes a settings conversation instead of a re-integration project.
Whole-corpus delivery is the norm here rather than the exception: roughly 725 of the 1,744 records in Datadory's catalog arrive as complete file sets rather than per-series requests. The difference is who babysits the loop.
Which petroleum series matter for storage and transportation desks?
Weekly inventories lead the list. A representative series, PET.WO6ST_R20_1.W, tracks Midwest (PADD 2) ending stocks of conventional CBOB gasoline blending components in thousand barrels; its history in the current capture runs from April 9, 2004 to January 16, 2026, with the agency's most recent revision stamped January 22, 2026 - about a week between an observation week and the final print, which is the lag to budget for when quoting the freshest week.
Its three newest observations read 40,412, 38,357 and 35,617 thousand barrels for the weeks ending January 16, January 9 and January 2, 2026 - a 4.8-million-barrel two-week build visible in three lines of text. That structure generalizes across the file: weekly PADD, lower-48 and state stocks plus import series carry observation histories observed from April 2004 forward, while the imports companion adds city, port and refinery detail segmented by crude grade and country of origin. Traders tracking inventory cycles and quants assembling positioning features end up reading the same lines.
What does one row look like?
One record per EIA series, exactly as the bulk text returns it - this shape was verified against the actual file during catalog research:
# One record per EIA series, exactly as the bulk text carries it
series_id PET.WO6ST_R20_1.W
name Midwest (PADD 2) Ending Stocks of Conventional CBOB
Gasoline Blending Components, Weekly
units Thousand Barrels
frequency W (weekly)
geography USA-IA+USA-IL+USA-IN+USA-KS+USA-KY+USA-MI+USA-MN+
USA-MO+USA-ND+USA-NE+USA-OH+USA-OK+USA-SD+USA-TN+USA-WI
start 20040409 end 20260116
last_updated 2026-01-22T21:12:35-05:00
data (latest) 20260116 40412 | 20260109 38357 | 20260102 35617Two things fall straight out of that shape. The geography spec enumerates all fifteen states summed into PADD 2 and rides the row itself, so a dashboard at district level and a state-by-state study share one parser - climbing or descending the geographic ladder changes no schema. And because the revision stamp travels with every series, vintage handling becomes a filter downstream rather than a footnote nobody reads.
How far does the history reach, and what are the edges?
- Geography: United States at national, PADD and state level for stocks; imports descend to city, port and refinery grain by crude grade and country of origin; the international sets extend coverage to individual countries.
- Temporal: weekly stocks observed from April 2004 forward, Annual Energy Outlook editions reaching back to 2014, and every version ending cleanly at its capture date - August 16, 2026 for 81.0.0.
- Granularity: individual EIA series, tens of thousands in the petroleum file alone, each carrying its complete observation history rather than a windowed excerpt.
Two edges deserve planning rather than surprise. First, revision lag: expect roughly a week between an observation week and the agency's final print, stamped in the record itself. Second, the capture edge: anything after a version's assembly date arrives in the next numbered version, which lands on your schedule once delivery is configured - never as a silent patch to history you already quoted.
Who builds on the EIA petroleum series?
Ranked by how directly the weekly, geographically-split grain answers the day job:
- Investors and quant researchers. Weekly PADD and state stocks are the supply-demand balance gauge behind inventory-cycle positioning; investors and quants use cases shows where the series slot into screens and backtests.
- Data scientists and ML engineers. Pin training inputs to a numbered version so a rebuild reproduces bit-for-bit; the pattern is laid out in data scientists use cases.
- Developers and builders. One scheduled versioned object replaces per-series request loops and pagination arithmetic - plumbing that stops being maintained by hand.
- Supply and trading analysts. Imports sliced by port, refinery, crude grade and origin turn origin-shift questions - which Gulf Coast grades came from where this quarter - into queries.
- Journalists and academics. Cite one named version and the figures stay checkable indefinitely, which is precisely the property citation-grade work depends on.
Persona tagging agrees: data scientists plus developers score four on relevance, journalists and academics two, investors and market researchers one.
Which datasets pair with the EIA volumes?
This archive answers how much; the rest of the shelf answers where. OGIM — Oil and Gas Infrastructure Mapping database (EDF) scores 9/10 and maps the physical estate: 6.7 million harmonized features across 152 countries - 4.54 million wells, 1.86 million pipeline segments spanning more than 1.2 million kilometres, 3,661 petroleum terminals, 692 refineries and 547 LNG facilities. The head-to-head comparison scores that volumes-versus-locations trade-off line by line, and the strongest desks run both.
Closer to the tank, HIFLD Petroleum Terminals maps all 2,302 operable US bulk terminals with fifty attributes apiece, TGS US Pipelines contributes 731,388 segments carrying operator, status, diameter, capacity and flow direction, OpenStreetMap adds the only realtime layer - roughly 867,000 tagged storage tanks - and TankTerminals.com goes commercial-global with capacity fact sheets across 13,000-plus facilities in 2,420 ports. Join key throughout: geography plus product plus period.
Where to go next
Start with the oil & gas storage & transportation data guide for the full eight-record pool, or the industry hub for how the volumes layer sits beside the asset layers. For the gas-side twin of this workflow, see how EIA weekly storage report history arrives as one continuous regional panel. The PUDL Raw EIA Bulk API Data dataset page documents sample rows, the field dictionary and coverage in full.
Request a sample scoped to the packages, series families and era you actually need - real rows come back with the manifest attached before anything recurring switches on.
| Record | Grain | History | What it adds |
|---|---|---|---|
| OGIM — Oil and Gas Infrastructure Mapping database (EDF) | Asset-level features across 152 countries | Curated compilation, version 2.7 | Wells, pipeline segments, terminals, refineries, LNG and compressor stations - the where behind the how much |
| HIFLD Petroleum Terminals (POL Terminals) | 2,302 operable US terminals at address level | Most recent edition revalidated August 2025 | Shell storage capacity in barrels, primary commodity and access-mode flags for terminal screening |
| TGS US Pipelines ArcGIS Feature Layer | 731,388 line segments, lower 48 plus Alaska | Attribute maintenance stamps per segment | Operator, status, diameter, capacity and flow direction along roughly 44 fields |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
PUDL Raw EIA Bulk API Data (archived snapshots)
None
OGIM — Oil and Gas Infrastructure Mapping database
HIFLD Petroleum Terminals (POL Terminals)
TGS US Pipelines ArcGIS Feature Layer
OpenStreetMap — Pipeline & Storage Facility Features
man_made · pipeline · substation …+15 more
TankTerminals.com — Global Tank Terminal Directory
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What exactly comes in the EIA bulk petroleum archive?
Every bulk series the U.S. Energy Information Administration publishes at one instant. Version 81.0.0, assembled August 16, 2026, carries 30 files near 1.6 GB: the petroleum and other liquid fuels package (57.5 MB), the petroleum imports package (5.5 MB) and the natural gas package (4.6 MB), alongside outlooks, coal, electricity, State Energy Data System and international sets.
How far back do the weekly petroleum stock histories go?
Weekly lower-48, PADD and state stocks plus import series carry observation histories observed from April 2004 forward. The worked example - Midwest PADD 2 ending stocks of conventional CBOB gasoline blending components - runs April 9, 2004 to January 16, 2026 in thousand barrels.
Will the numbers change after I build on them?
No. Each numbered version is immutable once published, and the manifest records size and hash for every resource, so a rebuild reproduces bit-for-bit. Newer captures arrive as separately numbered versions, making freshness a pinning decision instead of silent drift underneath your work.
Can delivery be scoped to just the storage-and-transportation packages?
Yes. Samples are scoped to the packages, series families and era you name, and most evaluations never touch the full 1.6 GB - the three logistics-relevant packages together weigh well under 70 MB. Rows land typed, with fill rates noted and the manifest attached.
Which datasets pair with the EIA time series?
Asset-side layers that answer where the volumes sit: OGIM's 6.7 million infrastructure features across 152 countries, HIFLD's 2,302 US terminals, TGS's 731,388 pipeline segments, OpenStreetMap's roughly 867,000 tagged storage tanks and TankTerminals.com's 13,000-plus facility directory.