Datadory notebook
Petroleum terminals shapefile: where fuel sits, delivered as rows
Datadory delivers petroleum terminal data covering the places fuel physically enters, leaves and sits: a national United States census of 2,302 operable bulk terminals, each pinned with location, owner, shell storage capacity in barrels, commodity class and twenty YES/NO access-and-product flags; 3,661 global terminal points across 152 countries sitting beside refineries, LNG facilities and 1.86 million pipeline segments in one schema; and operational fact sheets for 13,000-plus facilities across 2,420 ports with tank counts, berths, throughput and named managers - delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
What counts as petroleum terminal data?
The query sounds like a file format; the decision underneath it is a supply map. A petroleum terminal is where product changes custody - ship to shore, pipeline to rack, refinery to market - which makes terminal layers the skeleton every fuels analysis hangs on. Three different products travel under the one search term, and they settle different questions.
- An infrastructure census inventories facilities: one row per terminal, location resolved to coordinates, capacity in barrels, owner named, and the access modes and product slate spelled out flag by flag. It answers where can product physically enter or leave this market.
- A commercial directory goes past existence into operations: tank counts and size ranges, berth occupancy, throughput, turnaround times, and the people who run the site. It answers how hard does this facility actually work, and who answers its phone.
- A community map inventories individual vessels rather than facilities: hundreds of thousands of tagged storage tanks distinguished by content. It answers where is one specific tank when facility grain is still too coarse.
Datadory ships all three grains side by side under one delivery contract, which is what turns a pile of pin maps into a network you can query.
Which datasets cover the world's petroleum terminals?
Five named products cover this request, running from national depth to global frame to commercial edge. Figures come from Datadory's catalog as of August 2026.
- HIFLD Petroleum Terminals (POL Terminals) - the national census. Produced for the Department of Homeland Security by Oak Ridge National Laboratory with Argonne National Laboratory and the NGA Homeland Security Infrastructure Program team, it maps 2,302 operable US bulk petroleum terminals across all fifty states plus Puerto Rico, the US Virgin Islands, Guam and the Northern Mariana Islands, roughly fifty attributes per row. Its most recent revision revalidated 658 records, added ten terminals and removed forty-seven confirmed-closed, duplicated or merged facilities. Scored 7 out of 10.
- OGIM - Oil and Gas Infrastructure Mapping database (EDF) - the global frame. Environmental Defense Fund's census, assembled in support of the MethaneSAT mission, compiles 188 curated sources into roughly 6.7 million features across sixteen layers and 152 countries - including 3,661 petroleum terminals beside 692 crude oil refineries, 547 LNG facilities and 1.86 million pipeline segments spanning more than 1.2 million kilometres. Every feature carries a unique identifier, region, country and provenance reference. The highest-scoring record in the slice at 9 out of 10.
- TankTerminals.com - Global Tank Terminal Directory - the operating picture. 13,000-plus tank terminals and production facilities organized across 2,420 ports and cities, each fact sheet carrying total tank capacity, tank count and size range, berth count, stored products, modal connections and up to three named managers with direct contacts, extended by throughput, berth occupancy, vessel turnaround and tank-turn panels down to berth grain. Scored 5 out of 10.
- TGS US Pipelines ArcGIS Feature Layer - the connective tissue. 731,388 US oil, gas and refined-product polyline segments, roughly 44 attributes each: operator, status, nominal diameter, liquids capacity in Mbbl/d or gas capacity in MMcf/d, and flow direction. Terminals tell you where product can land; this tells you how it travels between them. Scored 7 out of 10.
- OpenStreetMap - Pipeline & Storage Facility Features - the vessel census. Roughly 867,000 features tagged as storage tanks worldwide, distinguishable by content - fuel, gas, LNG, LPG - mapped feature by feature by a contributor base that never stops editing. Scored 6 out of 10.
What fields does each terminal record carry?
The census's roughly fifty attributes beat a coordinates-only gazetteer because they qualify every facility instead of merely placing it. The dictionary spine:
- Identity and placement - a stable terminal identifier patterned from a fixed prefix plus state FIPS plus sequence, the facility name, full street address through ZIP+4, county name and five-digit county FIPS, and point coordinates in decimal degrees.
- Classification - facility type and operating status as coded enums, typically BULK TERMINAL and IN SERVICE, with the NAICS code and description beside them - 424710, petroleum bulk stations and terminals, on essentially every row.
- Capacity - total bulk shell storage in barrels, the number behind every capacity league chart.
- Ownership and commodity - owner and operator names where reported, plus a primary commodity class from a standardized domain with comma-list extensions such as REFINED, BIOFUEL.
- Provenance - originating agency, source date, validation method and validation date on every row.
The flag blocks are what make the layer analytical rather than cartographic. Eight YES/NO pairs spell out inbound and outbound access by truck, pipeline, marine and rail - the inclusion rule itself is visible in them, since receiving product by tanker, barge or pipeline qualifies a terminal on its own. Twelve further product-handling flags cover asphalt through avgas, so a gasoline-only network, a crude-capable subset or an ethanol-blending corridor is a filter rather than a research project. And the provenance fields turn vintage into a column read: attribute dates cluster in particular survey years, so aging any record before relying on it takes one query instead of an archaeology dig.
What do delivered rows look like?
Captured during the August 2026 research pass. Four Pennsylvania rows - a state that packs its pipeline corridors with terminals - show the anatomy every record shares:
TERM_ID NAME CITY STATE TYPE STATUS COMMODITY CAPACITY_barrels
ANLTK42001 LUCKNOW-HIGHSPIRE TERMINAL - ALLENTOWN ALLENTOWN PA BULK TERMINAL IN SERVICE REFINED 150100
ANLTK42002 ZENITH ENERGY HOLDINGS - ALTOONA TERMINAL ALTOONA PA BULK TERMINAL IN SERVICE REFINED 163000
ANLTK42003 GULF OIL, LP ALTOONA TERMINAL ALTOONA PA BULK TERMINAL IN SERVICE REFINED 1440000
ANLTK42004 SUNOCO PARTNERS MKTG & TERMINALS - ALTOONA ALTOONA PA BULK TERMINAL IN SERVICE REFINED, BIOFUEL 95748Read the anatomy rather than the digits. One row per terminal, keyed by an identifier patterned from a fixed prefix plus state FIPS plus sequence - the join key across every extract you will ever cut. Capacity is total bulk shell storage in barrels, and it spreads further inside one town than across most regions: Gulf Oil's Altoona farm holds 1,440,000 barrels while Sunoco's terminal across the same ZIP code holds 95,748, an order-of-magnitude difference that survives any regional average. Commodity arrives as a primary class with comma-list extensions where a site handles several - the biofuel flag on a refined-products terminal is exactly the kind of fact a blending study needs and a pin map omits.
How wide does terminal coverage run?
Geography splits by product, deliberately. The census is genuinely national - Alaska's terminals and Hawaii's rack network sit under the same schema as Gulf Coast marine complexes, and territory fuel logistics in Puerto Rico, the US Virgin Islands, Guam and the Northern Marianas is modeled rather than omitted. The global mapping database crosses every border the census stops at: 152 countries and six continents under one attribute schema, with North America and Europe running dense and some national contributions thinner. The directory reaches the Asian storage buildout and Middle East export complexes that national inventories never cross, organized port by port across 2,420 cities.
Temporal behaves like an asset register, not a ticker. Census editions arrive as snapshots, and each row carries its own source and validation stamps - attribute vintages cluster in the mid-2010s survey years, so screening questions tolerate the age happily while anything time-sensitive needs a companion series. In Datadory deliveries every extract states the release it was cut from, and revisions arrive as new vintages beside the old rather than overwriting them, so a map built last quarter stays reproducible.
Granularity is the facility itself: one point per terminal with address-level attributes in the census, one fact sheet per facility in the directory, one tagged node or way per tank in the community map. None of the three rolls up to market aggregates - which is precisely why any market aggregate you want, you can compute honestly.
How do you build a terminal base map that holds together?
The standard move: load the census for the United States, the global mapping layer for everywhere else, and connect both through the segment-level pipeline inventory where diameter, capacity and flow direction matter. Three details decide whether the merged workspace behaves.
- Geometry differs by design. The census ships point and polygon classes for facilities while the global layer's midstream features are points; decide per question whether a terminal is a dot or a footprint before styling anything.
- Identifiers never match across sources. Each product keys rows in its own universe - a census terminal ID and a global feature ID share no prefix, no format and no history. Join on normalized names plus proximity, keep both keys afterward, and treat a fuzzy match below your distance threshold as unmatched rather than heroic.
- Vintage varies by row, not by file. Source stamps differ feature by feature, running older in some countries than others. Filter or annotate by the vintage column before publishing any count.
Then give the map a pulse. Frozen archive snapshots of the federal inventory series put weekly PADD and state product stocks beside the static assets - the pairing behind most working storage models - and the head-to-head between the two records works through it cell by cell in OGIM vs PUDL Raw EIA Bulk API Data. When freshness beats completeness, the community map earns its place: it is the only continuously edited layer in the set, so a count taken today describes today.
Which terminal dataset fits which project?
Read the table as a routing decision, not a ranking. Breadth, depth and operating detail trade against each other cleanly, and most production workflows end up joining two of the five.
- US market entry, fuel logistics, emergency planning - the census. Fifty attributes per named terminal beat every alternative when the question is which American facilities exist, how big they are and how product reaches them.
- Global screening under one schema - the mapping database. A 152-country query beats 152 integrations, and terminals arrive beside refineries, LNG facilities and pipeline segments, so corridor context needs no second source.
- Commercial diligence on a specific facility - the directory. Tank counts, berth occupancy, throughput and named managers answer questions an infrastructure census is structurally silent on.
- Network and routing work - the segment inventory, paired with whichever terminal layer frames the study; the head-to-head in HIFLD Petroleum Terminals vs TGS US Pipelines works through that exact pairing.
- Vessel-grain questions - the community map, where roughly 867,000 individually tagged tanks settle what one specific tank holds even when no facility record mentions it.
Who builds on petroleum terminal data?
Ranked by how directly the layers answer the day job:
- Fuel logistics and network planners model truck, rail, barge and pipeline catchments around storage - the eight access-mode flags turn inbound and outbound mode into a filter instead of a guess, and shell capacity sizes the node. Start at developers builders use cases.
- Investors and quant researchers read midstream and downstream concentration off terminal density and capacity rather than company decks, pairing asset footprints with weekly inventory series for the flow leg. See investors quants use cases.
- Market researchers and consultants size storage markets port by port - the directory's capacity-evolution panel runs forward as well as backward, so entry screens can price buildouts that have not finished. See market researchers use cases.
- Emergency management and continuity teams keep pre-event inventories of where fuel physically sits, status-flagged and county-coded so regional rollups assemble without geocoding work.
- Data scientists and ML engineers inherit labeled spatial covariates at national scale - access flags and commodity classes make clean features, and stable per-terminal keys make panel construction boring in the best way. See data scientists use cases.
How is petroleum terminal data delivered through Datadory?
API, files, or your warehouse. Daily, weekly, or hourly.
Scoping is the point. Want only marine-served terminals above half a million barrels on the Gulf Coast? Every jet-fuel-capable site inside a three-state corridor? Directory fact sheets for one port range, refreshed weekly while the base layer runs monthly? Name the filters when you request a sample and the sample arrives already cut to match, with the production feed following the identical shape - anything prototyped on the sample survives delivery intact.
Why get terminal data through Datadory?
Because the terminals were never the hard part - the seams between them are. Identifiers that share no universe, forcing name-plus-proximity joins nobody documents. Capacity columns that measure shell geometry rather than throughput, tempting every first-time user into a utilization claim the field cannot support. Unreported contact fields carrying an explicit placeholder convention that collides with CRM tables expecting empty strings. Snapshot vintages hiding inside a file whose header claims one year. Each seam is survivable once; together they are why two terminal dashboards disagree.
Datadory handles the seams upstream of you: one typed schema across sources, keys preserved and cross-walked, capacity semantics documented beside every value, placeholders normalized, and each extract stamped with the release it came from. Say which states, commodities and capacity bands you model when you request a sample and it arrives shaped to that scope - the feed that follows matches it exactly.
Where to go next
Start with the HIFLD Petroleum Terminals (POL Terminals) dataset page for the full fifty-field dictionary and to request a sample cut to your states, then the OGIM - Oil and Gas Infrastructure Mapping database (EDF) page for the global frame and its 188-source provenance catalog. For the network between the terminals, read TGS US Pipelines ArcGIS Feature Layer and its head-to-head in HIFLD Petroleum Terminals vs TGS US Pipelines; for operating depth, the TankTerminals.com - Global Tank Terminal Directory page covers throughput, berths and named contacts.
The oil & gas storage & transportation data hub holds all eight pooled records on one page, best oil gas storage transportation datasets ranks them, and the oil gas storage and transportation data guide maps the whole slice. Adjacent deep-dives: the global oil pipeline database GeoJSON guide for routable lines, and the EIA weekly storage report history guide for the weekly numbers that move through the pipes.
| Field | Type | Definition | Example |
|---|---|---|---|
| TERM_ID | string | Unique terminal identifier, patterned ANLTK plus state FIPS plus sequence; the join key across extracts. | ANLTK42001 |
| NAME | string | Terminal or company facility name. | LUCKNOW-HIGHSPIRE TERMINAL - ALLENTOWN |
| ADDRESS / CITY / STATE / ZIP / ZIP4 | string | Physical street-address components of the terminal; unreported values carry an explicit placeholder rather than nulls. | ALLENTOWN, PA 18109 |
| TYPE | enum | Facility type drawn from a standardized coded domain. | BULK TERMINAL |
| STATUS | enum | Operating status of the terminal. | IN SERVICE |
| COUNTY / COUNTYFIPS | string | County name and five-digit county FIPS code, so regional rollups assemble without geocoding. | LEHIGH / 42077 |
| LATITUDE / LONGITUDE | geo | Point coordinates in decimal degrees. | 40.63078, -75.431764 |
| NAICS_CODE / NAICS_DESC | string | Facility industry classification, typically petroleum bulk stations and terminals. | 424710 PETROLEUM BULK STATIONS AND TERMINALS |
| OWNER / OPERATOR | string | Terminal owner and operating company where reported. | ZENITH ENERGY US LP |
| COMMODITY | enum | Primary commodity class handled, from a standardized coded domain; comma lists allowed. | REFINED, BIOFUEL |
| CAPACITY | number | Total bulk shell storage capacity in barrels - tank geometry, not annual throughput. | 1440000 |
| TRUCK_IN/OUT, PIPE_IN/OUT, MARINE_IN/OUT, RAIL_IN/OUT | boolean x8 | YES/NO inbound and outbound flag for each transport mode serving the terminal. | TRUCK_IN=YES, PIPE_IN=YES, MARINE_IN=NO |
| ASPHALT through AVGAS | boolean x12 | Product-handling flags indicating which commodity groups the terminal stores or moves. | GASOLINE=YES, DISTILLATE=YES, CRUDE_OIL=NO |
| SOURCE / SOURCEDATE / VAL_METHOD / VAL_DATE | string | Originating agency and publication date plus geometry-validation method and date; the vintage read. | EIA / 20150727 / IMAGERY |
| Dimension | Coverage |
|---|---|
| Geography | United States - all fifty states plus Puerto Rico, the US Virgin Islands, Guam and the Northern Mariana Islands; a genuine national file rather than a lower-48 cut |
| Temporal | Snapshot editions tracked as an asset register; every row carries its own source-agency, source-date, validation-method and validation-date stamps, and delivered extracts state the release they were cut from |
| Granularity | One point per terminal with address-level attributes - facility resolution, not market aggregates; polygons available where footprint matters |
| Qualification rule | A terminal enters with 50,000 barrels or more of bulk shell storage, or receipt capability from tanker, barge or pipeline; closed and duplicate facilities are removed in revision passes |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
HIFLD Petroleum Terminals (POL Terminals)
OGIM — Oil and Gas Infrastructure Mapping database
TankTerminals.com — Global Tank Terminal Directory
TGS US Pipelines ArcGIS Feature Layer
OpenStreetMap — Pipeline & Storage Facility Features
man_made · pipeline · substation …+15 more
PUDL Raw EIA Bulk API Data (archived snapshots)
None
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What qualifies a facility to appear as a petroleum terminal?
In the United States census, two routes: the terminal holds 50,000 barrels or more of total bulk shell storage capacity, or it can receive product from tanker, barge or pipeline regardless of shell size. Records must be operable, the standard classification is NAICS 424710 - petroleum bulk stations and terminals - and revision passes remove confirmed closures, duplicates and sites merged into neighbors.
Does terminal data include storage capacity and commodity detail?
Yes. Each census row carries total bulk shell storage in barrels - Pennsylvania samples span 95,748 barrels at one Altoona terminal to 1,440,000 barrels at another in the same town - plus a primary commodity class such as REFINED or BIOFUEL and twelve product-handling flags covering asphalt, chemicals, propane, butane, refined products, ethanol, biodiesel, crude oil, jet fuel, gasoline, distillate and avgas.
Is there global terminal coverage beyond the United States?
Two ways. The Oil and Gas Infrastructure Mapping database resolves 3,661 petroleum terminals across 152 countries on one schema beside refineries, LNG facilities and pipeline segments, every feature keyed by a unique identifier with country and region attached. The Global Tank Terminal Directory trades breadth for operating depth: 13,000-plus facilities across 2,420 ports and cities with tank counts, berth occupancy, throughput estimates and up to three named managers per fact sheet.
Which datasets pair with terminal data?
Three carry the surrounding work: the US pipelines segment inventory connects terminals into a routable network with diameter, capacity and flow direction; the archived energy-administration snapshots add weekly PADD and state inventory levels back to April 2004 for the volumes leg; and the storage-tank mapping layer covers individual vessels when the question is a single tank rather than a facility.