PUDL Raw EIA Bulk API Data (archived snapshots)
Datadory delivers pudl raw eia bulk api data archived snapshots covering every bulk series the U.S. Energy Information Administration publishes - weekly PADD and state petroleum stocks, imports tracked by port, refinery and crude grade, and natural gas statistics - frozen in named versions like 81.0.0 (August 16, 2026), thirty files near 1.6 GB strong, on your cadence.
What is PUDL Raw EIA Bulk API Data (archived snapshots)?
PUDL Raw EIA Bulk API Data (archived snapshots) is the official U.S. energy-statistics surface held perfectly still, version by version. Catalyst Cooperative's Public Utility Data Liberation project gathers every bulk file the U.S. Energy Information Administration publishes and repackages each capture as numbered archive releases with a Frictionless datapackage.json manifest recording byte size and hash for every resource. Version 81.0.0, assembled August 16, 2026, carries 30 files near 1.6 GB. For storage and transportation work, three of those files do almost all the lifting: the petroleum and other liquid fuels package (57.5 MB) spanning production, imports, refining, exports, prices, consumption, stocks and reserves; the petroleum imports package (5.5 MB) detailing movements at national, PADD, state, city, port and refinery level by crude grade and country of origin; and the natural gas package (4.6 MB) with production, pipelines, exports, prices, consumption, stocks and reserves.
Everything else rides along on the same record - Annual Energy Outlook editions from 2014 through 2025, the International Energy Outlook, the Short-Term Energy Outlook, coal statistics down to mine level, electric system operating data, plant-level generation and fuel data, CO2 emissions aggregates, State Energy Data System tables, nuclear outages and international series. Unpacked, each package yields the raw EIA bulk text - PET.txt for petroleum - holding one JSON object per data series, each with its complete observation history embedded.
The pin is the point. Where infrastructure maps tell you where terminals and pipelines sit, this corpus tells you how much product sits and moves there, and because every figure traces to a named, hashed version, a number cited today still checks out years from now. It sits in the oil gas storage transportation data hub as the slice's only deep time-series play. Get a sample of this dataset before anything else.
What do sample rows look like?
One record per EIA series, exactly as it reads out of the bulk text - this one was checked against the actual file at research time:
series_id PET.WO6ST_R20_1.W
name Midwest (PADD 2) Ending Stocks of Conventional CBOB
Gasoline Blending Components, Weekly
units Thousand Barrels
frequency W (weekly)
geography USA-IA+USA-IL+USA-IN+USA-KS+USA-KY+USA-MI+USA-MN+
USA-MO+USA-ND+USA-NE+USA-OH+USA-OK+USA-SD+USA-TN+USA-WI
start 20040409
end 20260116
data (latest) 20260116 40412 | 20260109 38357 | 20260102 35617Read it apart and the structure shows itself: the geography spec enumerates all fifteen states summed into PADD 2, the observation dates run on EIA's Friday week-ending convention, and the values descend through January 2026 in thousand-barrel terms - 40,412, then 38,357, then 35,617. The series opens April 9, 2004, so one line carries two decades of weekly inventory history.
The three logistics-weighted packages inside version 81.0.0:
package size
petroleum-and-other-liquid-fuels 57.5 MB
petroleum-imports 5.5 MB
natural-gas 4.6 MBThose three together weigh well under 70 MB of the roughly 1.6 GB deposit, which is why a scoped sample rarely needs the full payload. Your sample arrives carrying the series families and era you actually name.
What fields does the dataset include?
Every series line speaks the same eight-field dialect, verified against the bulk text rather than inferred from documentation:
A few riders sit outside the headline dictionary: each series also carries EIA's own prose description, an attribution string naming the originating agency, and a copyright notice field that reads None on the overwhelming majority of lines. Whether the natural gas package itemizes dedicated underground-storage series beyond its general stocks lines was not itemized at research time - those get spelled out against your sample rather than guessed at here, and additional fields are available on request.
What does coverage look like across geography, time and granularity?
- Geo: United States at national, PADD, state, city, port and refinery level depending on the series; international energy datasets extend coverage to individual countries
- Temporal: weekly petroleum stocks observed from April 2004 forward, with some series running longer and Annual Energy Outlook editions reaching back to 2014; each version is a frozen capture ending at its assembly date
- Granularity: individual EIA data series - tens of thousands in the petroleum file alone - each carrying its complete observation history rather than a windowed excerpt
Breadth over single-program depth is the trade. One corpus spans every EIA program where rivals go deeper on a single table family, so a fuels-desk analyst and a power-market analyst can pull from the same object. Because geography and periods are inherited from whatever the agency publishes, coverage quirks - a series that stops, a geographic cut that only exists above state level - are part of the record, auditable in the manifest, instead of surprises discovered mid-project.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Pick the packages, pick the cadence, pick the landing zone. The version pin travels with every extract, so provenance survives the trip: a dashboard, a model's training table and a board deck can all cite the same numbered capture. Weekly, monthly, annual and daily renditions arrive together where a program publishes them, rather than as separate integrations to babysit. Because each version ends at its capture date, staying past that edge is a scheduling decision handled on our side - newer captures slot in as their own numbered versions, never as edits underneath you. The sample comes first: name the packages, the series families and the era, and it lands shaped to that scope with the manifest attached. See how builders wire frozen corpora into scheduled pipelines in developers builders use cases.
Who uses this data, and for what?
- Reproducible pipelines and model training - pin training inputs to a numbered version so a rebuild reproduces bit-for-bit; the pattern is laid out in data scientists use cases
- Scheduled delivery plumbing - replace per-series request loops and pagination arithmetic with one versioned object that lands on a schedule
- Inventory-cycle signals - weekly PADD and state stocks as the supply-demand balance gauge behind positioning work; investors quants use cases shows where the series slot in
- Trade-flow forensics - imports sliced by port, refinery, crude grade and country of origin for origin-shift studies
- Citation-grade research and journalism - cite one named version and the figures stay checkable indefinitely, the property archival work depends on
Persona tagging puts data scientists and ML engineers plus developers and builders at the top (relevance four each), journalists and academics next at two, investors and quants and market researchers at one. Sales and growth teams, competitive intelligence and e-commerce operators score zero - aggregates without company-level detail give them nothing to act on. The shape tells you what this is: infrastructure underneath analyses, rarely the headline story itself.
Which notes and neighboring datasets pair with it?
The neighbors split cleanly along the volumes-versus-locations line. OGIM Oil and Gas Infrastructure Mapping Database (EDF) scores higher overall (9 versus 8) and maps 6.7 million infrastructure features across 152 countries, but carries no weekly volumes; the head-to-head comparison scores that trade-off line by line, and the strongest desks run both - asset layers for where, this archive for how much. HIFLD Petroleum Terminals (POL Terminals) maps 2,302 US terminals with fifty attributes apiece; TankTerminals.com Global Tank Terminal Directory goes global on terminal capacity; OpenStreetMap Pipeline & Storage Facility Features and the TGS US Pipelines ArcGIS Feature Layer round out the physical network. None of them move.
Glossary notes sharpen the vocabulary before you scope anything: what a static snapshot commits you to, how to judge dataset snapshot staleness per version, what PADD districts actually bound, why petroleum stocks inventories are read as a balance gauge, what makes a long-format time series pipeline-friendly, how this relates to the EIA weekly storage report tradition and what a gas storage inventory tracks. The publishing source profile lives at Catalyst Cooperative via Zenodo, and the best oil gas storage transportation datasets ranking shows where this record lands in the slice.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
series_id | string | EIA hierarchical series identifier encoding dataset family (PET, NG, SEDS and others), route, product, geography and frequency | PET.WO6ST_R20_1.W |
name | string | Human-readable series title describing measure, geography and frequency | Midwest (PADD 2) Ending Stocks of Conventional CBOB Gasoline Blending Components, Weekly |
units / unitsshort | string | Unit of measure in full and abbreviated form | Thousand Barrels / Mbbl |
f | enum | Frequency code of the series: A annual, Q quarterly, M monthly, W weekly, D daily | W |
geography | string | Geographic coverage expressed as an EIA geography spec, with state codes joined by + or - | USA-IA+USA-IL+...+USA-WI |
start / end | date | First and latest period present in the series, formatted YYYYMMDD | 20040409 / 20260116 |
last_updated | datetime | Timestamp of the agency's most recent revision to the series | 2026-01-22T21:12:35-05:00 |
data | array | Embedded time series as [period, value] pairs, newest first, one entry per observation week, month or year | ["20260116", 40412] |
Questions buyers ask
What exactly does version 81.0.0 freeze?
Everything the U.S. Energy Information Administration had published in bulk at one instant, August 16, 2026: 30 files near 1.6 GB, including the petroleum and other liquid fuels package (57.5 MB), the petroleum imports package (5.5 MB) and the natural gas package (4.6 MB), alongside coal, electricity, State Energy Data System, international and outlook resources.
How far back do the weekly petroleum stock histories reach?
The worked example runs weekly from April 9, 2004 to January 16, 2026 - two decades of Friday-dated PADD 2 gasoline blending component stocks. Other series stretch further, and Annual Energy Outlook editions reach back to 2014. Every version ends at its capture date rather than trailing off mid-history.
What geographic cuts do the petroleum import series support?
National, PADD, state, city, port and refinery levels, each broken out by crude grade and country of origin. That stack lets one corpus answer how much crude entered Gulf Coast refineries from a given origin and how much finished product sat in a specific district at week's end.
Will the numbers change after I build on them?
No. Each version is immutable once published, and the Frictionless manifest records byte size and hash for every resource, so a rebuild reproduces bit-for-bit. Newer captures arrive as separately numbered versions, making currency a pinning decision instead of silent drift underneath your work.
Can the corpus be cut down to just the storage-and-transportation files?
Yes. Samples are scoped to the packages and era you name, and most evaluations never touch the full 1.6 GB - the three logistics-relevant packages together weigh well under 70 MB compressed. Name the series families and period and the sample arrives shaped to that scope, manifest attached.
What remains unverified about the archive's internals?
Two gaps, flagged honestly: no schedule between published versions is documented anywhere, and whether the natural gas package itemizes dedicated underground-storage series beyond its general stocks lines was not confirmed at research time. Both are settled facts in a scoping sample, which is why the sample precedes any commitment.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.