PUDL Raw EIA Bulk API Archive
Datadory delivers pudl raw eia bulk api archive data covering every series the U.S. Energy Information Administration publishes in bulk, frozen as version 81.0.0 on August 16, 2026: natural gas, petroleum and liquid fuels, petroleum imports, coal, electricity, State Energy Data System tables and the full outlook shelf, across 29 packaged resources totaling about 1.6 GB.
What is the PUDL Raw EIA Bulk API Archive?
PUDL Raw EIA Bulk API Archive is the entire official U.S. energy-statistics surface pressed into one version-pinned object. Version 81.0.0, assembled by Catalyst Cooperative and published on August 16, 2026, captures everything the U.S. Energy Information Administration had released in bulk at that instant and repackages it as 29 ZIP resources plus a Frictionless datapackage.json manifest - about 1.6 GB in all. Bulk natural gas statistics, petroleum and other liquid fuels with supply, disposition, stocks and prices, petroleum import movements, coal production down to plant-level consumption, electricity generation, State Energy Data System tables, international energy series by country, CO2 emissions aggregates and carbon coefficients, nuclear outages and total energy series, and a full forecast shelf - Annual Energy Outlook editions from 2014 through 2026, International Energy Outlook editions from 2017 to current, and the Short-Term Energy Outlook - travel together in a single bundle.
The point of the pin is reproducibility. Against a live query interface, two analysts asking the same question a week apart can get different answers, and neither can prove afterwards which one was right. Here every number traces to a named version whose resources each declare a byte size, a stable hash and UTF-8 encoding in the manifest, so a result stays bit-for-bit identical years later. The archive began life as raw input for Catalyst Cooperative's Public Utility Data Liberation pipeline; its second career is as the cleanest way to hold a whole national energy corpus perfectly still. It sits in the oil gas exploration production data hub as the breadth play among single-program sources. Get a sample of this dataset before anything else.
What do sample rows look like?
One row per packaged resource - exactly how the manifest describes the payload:
resource format mediatype bytes
eiaapi-natural-gas.zip .zip application/zip 4611995
eiaapi-petroleum-and-other-liquid-fuels.zip .zip application/zip 57471282
eiaapi-petroleum-imports.zip .zip application/zip 5515297
eiaapi-coal.zip .zip application/zip 14292308
eiaapi-electricity.zip .zip application/zip 289894055
eiaapi-us-electric-system-operating-data-older-than-7d.zip .zip application/zip 682161345Byte counts above are captured from the version-81 manifest itself, not rounded marketing figures: the electric-system operating file alone accounts for 682,161,345 bytes, more than ten times the entire petroleum and liquid fuels package at 57,471,282, while bulk natural gas statistics weigh 4,611,995 bytes. Alongside these sit an 18.8 kB manifest and 22 further resources covering petroleum imports movements, State Energy Data System tables, international series, outlook editions and emissions aggregates. Every resource declares its format, mediatype, encoding and hash up front, so integrity can be established offline before a parser ever runs. Your sample arrives carrying the resources you actually need rather than the full gigabyte.
What fields does the dataset include?
The dictionary is the manifest's own vocabulary: one entry per packaged resource, each declaring name, path, format, size, encoding, hash and the EIA program it belongs to. The eight headline entries:
Resources beyond these eight headline entries fold into the same manifest structure and are spelled out against your sample rather than padded in here: the coal package with plant-level consumption, both electricity payloads including the 682,161,345-byte electric-system operating file, CO2 emissions aggregates and carbon coefficients, nuclear outage and total energy series, and the remaining Annual Energy Outlook and International Energy Outlook edition files. Each is described identically - name, path, format, bytes, hash, encoding and program tag - so nothing about the folded half changes shape, only how much of it your sample unpacks. Additional fields are available on request.
What does coverage look like across geography, time and granularity?
- Geo: United States at national, state and PADD level, plus country-level international energy series
- Temporal: one pinned instant - version 81.0.0 froze the corpus on August 16, 2026 - with underlying annual histories reaching back decades and Annual Energy Outlook projections running decades forward
- Granularity: national, state and country aggregates; monthly and annual periods depending on the program
Breadth over depth is the trade. Twenty-nine resources spanning every EIA program stand against single-program sources that go far deeper on one table family. Because geography and period are inherited from whatever the agency publishes, coverage quirks - a series that stops, a country that enters late, a forecast vintage that supersedes its predecessor - are part of the record rather than surprises discovered mid-project. The manifest makes each quirk auditable instead of anecdotal.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Pick the programs, pick the cadence, pick the landing zone. The version pin rides along with every extract, so provenance survives the trip into your warehouse: a dashboard, a model's training table and a client deck can all cite the same numbered snapshot. Monthly and annual renditions arrive together rather than as separate integrations. The sample comes first - name the resources and the era you need and it lands shaped to that scope, manifest attached, before any commitment. See how builders wire versioned corpora in developers builders use cases.
Who uses this data, and for what?
- Reproducible pipelines and model training - pin training inputs to a version so a rebuild reproduces bit-for-bit; the pattern is laid out in data scientists use cases
- Citation-grade research and journalism - cite a named version whose figures cannot drift under you, the property citation-grade work depends on
- Macro and quant factor construction - national, state, PADD and country-level aggregates in one pull instead of a dozen program-specific pulls; investors quants use cases shows where factor inputs slot in
- Client deliverables and benchmarking - the whole official-energy shelf behind a single engagement, scoped by program; see market researchers use cases
- Pipeline plumbing without pagination - replace thousands of live calls and their rate-limit arithmetic with one versioned object
In persona tagging this is plumbing with a three-star audience: data scientists and ML engineers rank highest, with investors and quants, developers and builders, market researchers, and journalists and academics all tagged at relevance two. The shape tells you what it is for - a reproducibility layer underneath analyses, rarely the headline analytical source itself.
Which notes and neighboring datasets pair with it?
The neighbors split by what they trade away. EIA International Energy Statistics is fresher query by query but carries no version pin, so a result can never be reproduced exactly. JODI Oil World Database and JODI Gas World Database standardize monthly supply and demand across participating economies, where this archive mirrors whatever one agency republishes in bulk. Global Registry of Fossil Fuels goes asset-level on reserves and emissions rather than aggregate statistics. Railroad Commission of Texas Data Sets (Production & Wellbore) is the closest peer by delivery model - both bulk objects - but goes lease-month deep on Texas wells where this trades depth for nationwide breadth; the head-to-head comparison scores the trade-off line by line. Outside the United States, NSTA UKCS Open Data plays the equivalent role for UKCS wells, production and infrastructure.
Three glossary notes sharpen the vocabulary before you request anything: what a static snapshot commits you to, why dataset snapshot staleness has to be judged per version, and how a bulk datapackage download differs from a pile of loose files. The publishing source profile lives at Catalyst Cooperative via Zenodo, and the best oil gas exploration production datasets ranking shows where this record lands in the slice.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
datapackage.json | text | Frictionless Data Package manifest listing all 29 resources with name, path, format, bytes, hash, encoding and the EIA data_set each belongs to | "name": "eiaapi-coal.zip", "bytes": 14292308 |
eiaapi-natural-gas.zip | text | Bulk EIA natural gas statistics series | 4,611,995 bytes |
eiaapi-petroleum-and-other-liquid-fuels.zip | text | Bulk EIA petroleum and other liquid fuels statistics, including supply, disposition, stocks and prices | 57,471,282 bytes |
eiaapi-petroleum-imports.zip | text | Bulk EIA petroleum imports movement data | 5,515,297 bytes |
eiaapi-international-energy-data.zip | text | International energy system time series by country | 24,043,433 bytes |
eiaapi-state-energy-data-system-seds.zip | text | State Energy Data System state-level production and consumption estimates | 9,467,055 bytes |
eiaapi-shortterm-energy-outlook.zip | text | Short-Term Energy Outlook forecast series | 5,421,504 bytes |
eiaapi-annual-energy-outlook-2026.zip | text | Annual Energy Outlook 2026 projection tables; editions 2014-2023, 2025 and 2026 ship as companion resources | 26,937,658 bytes |
Questions buyers ask
What exactly does version 81.0.0 freeze?
Everything the U.S. Energy Information Administration had published in bulk at a single instant, August 16, 2026: natural gas, petroleum and other liquid fuels, petroleum imports, coal, electricity, State Energy Data System tables, international series by country, emissions aggregates and the full outlook shelf, packaged as 29 ZIP resources plus a manifest totaling about 1.6 GB.
How large is the full archive, and can a sample be cut smaller?
About 1.6 GB across 29 ZIP resources plus an 18.8 kB manifest. The largest single resource, the electric-system operating file, is 682,161,345 bytes, while bulk natural gas statistics weigh 4,611,995 bytes. Samples are cut to the resources and era you name, so most evaluations never touch the full payload.
Will the numbers change after I build on them?
No. Each version is immutable once published: the file set, sizes and hashes never move, which is what makes citations and model rebuilds reproducible. Newer snapshots exist as their own numbered versions, so staying current is a pinning-and-delivery decision rather than a silent drift underneath your work.
Can I verify that what arrived matches the original?
Yes, and offline. The datapackage.json manifest records a hash, byte size, mediatype and UTF-8 encoding for every one of the 29 resources, so any copy can be checked against its declared digest before a parser runs. Integrity checking is a documented field of the product, not an afterthought.
Which outlook editions are included?
Annual Energy Outlook editions from 2014 through 2026, International Energy Outlook editions from 2017 to the current release, and the Short-Term Energy Outlook, each packaged as its own resource with the AEO2026 tables alone running 26,937,658 bytes. Projection vintages sit beside the historical series they once forecast, which is what makes forecast-error studies possible.
What remains unverified about the archive's internals?
Two documentation gaps, flagged honestly: the cadence between published versions is not stated anywhere, and whether the ZIPs unpack to line-delimited records or tables could not be confirmed without opening multi-hundred-megabyte files. Both are settled facts in a scoping sample, which is why the sample precedes any commitment.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.