GOGI — Global Oil & Gas Infrastructure Features Database
Datadory delivers GOGI Global Oil & Gas Infrastructure Features Database data covering the world's mapped oil and gas estate: more than 4.8 million features integrated from over 380 source datasets — wells with an active-wells cut, offshore platforms and onshore well pads, underground storage, refineries, stations, pipelines, ports and railways — organized into regional catalogs spanning every inhabited continent plus marine waters, delivered as analysis-ready spatial layers.
What is the GOGI Global Oil & Gas Infrastructure Features Database?
GOGI is a worldwide census of oil and gas physical infrastructure: more than 4.8 million features, integrated from over 380 source datasets into a single set of GIS-ready feature classes. It was produced by the US Department of Energy's National Energy Technology Laboratory (NETL) with the Environmental Defense Fund (EDF) under a cooperative research agreement, funded through the Climate and Clean Air Coalition Oil and Gas Methane Science Studies managed by UNEP — the program behind the global methane studies. Where a well registry answers "where was drilled," GOGI answers "what exists": the platforms, well pads, pipelines, refineries, storage, ports and railways that surround the well.
Eight feature classes organize the census. Under Production_Extraction sit Wells (including an active-wells subset), Platforms_and_Well_Pads and Underground_Storage. Facilities_Installations carries Refineries and Stations. Transport lays Pipelines and Railways as line features and Ports as points. Distribution follows geography rather than theme: a global catalog plus nine regional catalogs — North America, South America, Europe, Africa, Asia, Middle East, Australia, Antarctica — and a Marine set, with a Validation catalog alongside.
In Datadory's catalog of 1,744 datasets across 159 viable industries, this record scores 8/10, and it anchors the oil and gas drilling slice precisely because it breaks the slice's pattern. The other twenty primary records track rigs, permits, wells and production jurisdiction by jurisdiction; GOGI answers worldwide in one consistent geometry scheme, which makes it the base layer everything else pins onto.
Get a sample of this dataset
What does a sample of GOGI data look like?
Four feature families cover the census, plus a provenance stamp that rides on every feature regardless of family. Their anatomy, with values left as placeholders:
FAMILY: well point
feature_class : Wells
geometry : point
region_catalog : North America | South America | Europe | Africa | Asia
| Middle East | Australia | Antarctica | Marine | Global
status_cut : all compiled wells | Active Wells subset
FAMILY: platform or well pad
feature_class : Platforms_and_Well_Pads
geometry : point | polygon
setting : offshore platform | onshore pad
region_catalog : <regional or Marine>
FAMILY: transport line
feature_class : Pipelines | Railways
geometry : line
region_catalog : <region>
FAMILY: facility point
feature_class : Refineries | Stations | Underground_Storage | Ports
geometry : point
region_catalog : <region>
STAMP: carried by every feature
upstream_source: <which of the 380+ integrated datasets supplied the feature>The placeholders are deliberate: the documented structure names the feature classes, the geometries and the regional organization, not individual cell values, and Datadory does not print coordinates it cannot stand behind. Two things are already visible from the shapes alone. Every family keys on the same region-catalog dimension, so facilities stack against pipelines and wells in one map frame without reprojection work in between. And the upstream-source stamp travels with each feature, which turns the known overlap between merged inputs into an auditable deduplication pass rather than guesswork. Request a sample with your regions and feature classes named and the same five shapes come back filled.
What fields does the GOGI Features Database include?
Eight feature classes carry the database, mapped from the geodatabase's own catalog structure:
- Wells — point features compiled from global open sources, with an Active Wells subset layered on in the 2022 revision cycle.
- Platforms_and_Well_Pads — offshore platforms and onshore well pads, the surface expression of extraction on both settings.
- Underground_Storage — subsurface storage locations under the production-and-extraction grouping.
- Refineries — downstream conversion sites, mapped as points.
- Stations — processing, compressor and other midstream station points.
- Pipelines — line features tracing oil and gas gathering and transmission routes.
- Ports — maritime logistics points relevant to hydrocarbon movements.
- Railways — line features carrying the overland leg of crude and product logistics.
Attribute-level columns are a different story. Each of the 380+ upstream datasets contributes its own column vocabulary, and no single published schema reconciles them — so per-feature attributes beyond class membership (status codes, operator names, capacity figures, installation dates) sit under additional fields on request: named at sampling, confirmed on delivery, never guessed here. What the record guarantees outright is the geometry, the feature class, the regional catalog placement and the upstream-source stamp on all 4.8 million-plus features.
What does coverage look like across geography, time and granularity?
Geography - worldwide, organized for retrieval rather than for reading: a global catalog plus nine regional catalogs covering North America, South America, Europe, Africa, Asia, the Middle East, Australia and Antarctica, plus a dedicated Marine set for offshore features and a Validation catalog used during assembly.
Temporal - the compilation started life as the 2018 base inventory and advanced release by release: v1.0.2 arrived in February 2020, the Active Wells subset joined in March 2022, v1.0.3 repaired the Wells layer that April, and v1.0.3.1 shipped in June 2023 as the current release. Feature vintage therefore varies by region and class — which is exactly why the upstream-source stamp matters when you need to age a layer before relying on it.
Granularity - individual infrastructure features: discrete points for wells, platforms, pads, facilities and ports; discrete lines for pipelines and railways. This is asset-level resolution, not grid-cell aggregation, and each feature carries its attribution back to the specific source dataset it came from.
How is this dataset delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Name the regions and the feature classes when you request the sample — Gulf Coast refineries with connecting pipelines, North Sea platforms with the Marine catalog around them, or the full global census as it stands. The sample ships first either way; the ongoing feed lands on whatever cadence your models need.
Who uses this infrastructure data, and for what?
A global asset census earns its keep in four specific jobs:
- Data scientists and ML engineers — ready-made spatial covariates and labels. Distance-to-refinery, density-of-pads, proximity-to-transport become computable features instead of a data-engineering project spanning forty jurisdictions.
- Market researchers and consultants — the siting, supply-chain and market-entry context layer. A global view of where processing, storage and transport assets actually sit turns a market-sizing deck from assertion into map.
- Investors and quant researchers — infrastructure exposure mapping. Midstream and downstream footprints around producing basins support concentration analysis that well counts alone cannot deliver.
- Journalists, academics and students — provenance that survives review: a national-laboratory production with a documented methodology and a 594-page technical report (NETL Technical Report Series NETL-TRS-6-2018) behind the integration pipeline.
Developers and data-product builders get the fifth job: a clean global base layer to stand their own products on, without stitching regional sources together themselves.
Which personas get the most value?
Data scientists and ML engineers get the highest-leverage fit — 4.8 million-plus labeled spatial features from a single consistent scheme is training-set material, not a cleaning exercise. Journalists, academics and students get the citation-grade original, with the methods report answering the "how do you know" question before reviewers ask it. Market researchers and consultants get the context layer their drilling-focused alternatives lack: the assets around the well. Investors and quant researchers should pair it with something fresher — the June 2023 vintage caps signal recency, so treat GOGI as structural exposure rather than a timing signal. Start from the oil and gas drilling data hub, then see best oil and gas drilling datasets for where this census sits among regulator-first records.
Notes and related datasets
Provenance note — produced by the US Department of Energy's National Energy Technology Laboratory with the Environmental Defense Fund under a cooperative research agreement, funded through the Climate and Clean Air Coalition Oil and Gas Methane Science Studies managed by UNEP, with backing from major international oil and gas companies alongside EDF. One team, one integration pipeline, one geometry scheme.
Methodology note — the inventory was assembled with big-data search, custom data-integration algorithms and expert-driven curation across 380+ source datasets. Budget for geometry cleaning before spatial joins: overlapping inputs mean duplicate or near-duplicate geometries exist, and the geodatabase's Validation catalog relates to this step even though its precise role is not documented publicly. The methods report is the deep reference when a claim needs the pipeline behind the points.
Completeness note — Datadory scores this record 8/10. The feature-class dictionary above maps fully to the geodatabase catalog; attribute-level column vocabularies are confirmed per feature class at sampling under additional fields on request rather than asserted here, because they vary by upstream contributor.
Where to go next — pair the census with depth plays: EMODnet's human-activities layer adds annually revised European offshore well detail, while the OGIM infrastructure mapping database offers the other global feature pinning with versioned releases. The oil and gas drilling hub holds the regulator-first records that trade breadth for per-well attributes.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
Wells | point | Oil and gas well locations compiled from global open sources; includes an Active Wells subset layered on in March 2022. | Active Wells subset |
Platforms_and_Well_Pads | point | polygon | Offshore platforms and onshore well pads — the surface expression of extraction across both settings. | offshore platform, Marine catalog |
Underground_Storage | point | Subsurface storage locations grouped under production and extraction. | <storage site, <region> catalog> |
Refineries | point | Downstream conversion sites where crude is processed into products. | <refinery, <region> catalog> |
Stations | point | Processing, compressor and other midstream station locations. | compressor station, <region> catalog |
Pipelines | line | Oil and gas gathering and transmission routes traced as line geometry. | trunk line, <region> catalog |
Ports | point | Maritime logistics points relevant to hydrocarbon movements. | export terminal, Marine catalog |
Railways | line | Overland rail routes carrying the crude and product logistics leg. | rail corridor, <region> catalog |
Coverage — geography, temporal range, granularity
| Dimension | Coverage |
|---|---|
| Geography | Worldwide — a global catalog plus nine regional catalogs (North America, South America, Europe, Africa, Asia, Middle East, Australia, Antarctica), a Marine set, and a Validation catalog |
| Temporal | Base compilation published 2018; release lineage through v1.0.2 (Feb 2020), Active Wells addition (Mar 2022), v1.0.3 (Apr 2022) and current release v1.0.3.1 (June 2023); feature vintage varies by region and class |
| Granularity | Individual infrastructure features — discrete points and lines — each attributed to its upstream source dataset |
Questions buyers ask
What does the GOGI Global Oil & Gas Infrastructure Features Database include?
More than 4.8 million infrastructure features integrated from over 380 source datasets, organized into eight feature classes: Wells (with an Active Wells subset), Platforms_and_Well_Pads, Underground_Storage, Refineries, Stations, Pipelines, Ports and Railways. Content distributes across a global catalog, nine regional catalogs and a Marine set, with a Validation catalog from the assembly process.
Does GOGI cover offshore as well as onshore infrastructure?
Both, and it separates them cleanly. Offshore platforms land in Platforms_and_Well_Pads with the Marine catalog collecting offshore content, while onshore well pads occupy the same feature class on land. Ports and pipeline networks tie the two settings together logistically, so an offshore-to-onshore route reads as connected features rather than two unrelated layers.
Can you tell which wells are active?
Partially. An Active Wells subset was added to the Wells feature class in March 2022, giving a documented activity cut on top of the compiled global well points. Status vocabularies beyond that cut vary by upstream source, since each contributing dataset reports activity differently — deeper per-well attributes are confirmed under additional fields on request rather than promised blanket here.
How was the database assembled?
NETL and EDF combined big-data computing search, custom data-integration algorithms and expert-driven curation to identify and integrate more than 380 publicly documented source datasets worldwide. The full pipeline is documented in the 594-page NETL Technical Report Series NETL-TRS-6-2018, which covers source identification, harmonization decisions and validation — unusual methodological transparency for a global compilation.
How current are the features in GOGI?
The compilation began with the 2018 base inventory and progressed through v1.0.2 (February 2020), the Active Wells addition (March 2022), the Wells-layer repair v1.0.3 (April 2022) and the current release v1.0.3.1 (June 2023). Vintage varies by region and feature class, and each feature's upstream-source stamp tells you which dataset aged it — so currency checks happen per layer, not per marketing claim.
Who uses GOGI data?
Data scientists and ML engineers building spatial features and labels, market researchers mapping supply-chain and market-entry context, investors and quant researchers analyzing infrastructure exposure around producing basins, journalists and academics citing a national-laboratory compilation, and developers standing products on a global base layer without stitching dozens of regional sources.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.