Oil & Gas Equipment & Services · UK North Sea Transition Authority (NSTA)

NSTA National Data Repository (NDR)

Datadory delivers nsta national data repository ndr data as typed records keyed on Project IDs: the North Sea Transition Authority's statutory archive of UK Continental Shelf wells, seismic surveys, hazard surveys and reports, organised company-year-type-sequence from project level down to individual File IDs, with a completeness status riding on every item.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
United Kingdom Continental Shelf throughout - every project resolves against UK offshore Quadrant, Block and Field polygons, with hexagon density bins summarising where the archive runs deep
How far back
A continuous statutory archive rather than versioned releases: records span the depth of the UK offshore drilling era, each carrying its own completion dates in project metadata, and carbon dioxide appraisal and storage licence records join the shelf over time
How fine
Project-level entries - wells, seismic surveys, hazard surveys, reports - resolving down to individual File IDs, each tagged and statused

What is the NSTA National Data Repository (NDR)?

It is where the United Kingdom's offshore petroleum record lives by law rather than by choice. The North Sea Transition Authority (NSTA) operates the National Data Repository inside its Digital Energy Platform as the statutory archive for offshore petroleum-related licence information, with carbon dioxide appraisal and storage licence records being added over time alongside the conventional record.

Organisation is the archive's best idea. Every item hangs off a unique Project ID following the fixed pattern CCCCYYYYtypeNNNN - company, year, a project-type code (well = well description, seis = seismic survey, plus mhaz, rems and inpt), and an incrementing number. Decode one and you know who lodged the data, when, and roughly what it is before a single file opens. Files inside a project carry Classification Tags, and each information type carries one of three availability states - Available, Unavailable or Not Acquired - which makes the archive a completeness registry as much as a file store.

Scale sits at petabyte class, commonly cited around 1.6 PB - a figure that circulates widely even though the portal itself declines to publish one. Within Datadory's catalog of 3,932 datasets across 159 viable industries, this record scores 8/10: authority and depth rather than breadth. Get a sample of this dataset and judge the rows first.

What do sample rows from this dataset look like?

Start with the grammar, because that is what a statutory archive really sells: addressability. One Project ID, decoded:

# the anatomy - one Project ID groups everything lodged for a well, survey or report set
project_id       : CCCC2020seis0001
                   CCCC=company  YYYY=year  type=well|seis|mhaz|rems|inpt  NNNN=sequence

# fields riding on every project
file_id          : individual file handle within the project
c_tag            : Classification Tag - what kind of data item the file is
completeness     : Available | Unavailable | Not Acquired
geometry         : footprint against Quadrant, Block and Field polygons + hexagon density bins
project_metadata : date completed | main data type | licences | contractor | description | survey IDs

# the completeness registry, read as three states
Available    = file lodged and present
Unavailable  = acquired but never lodged
Not Acquired = never acquired at all

# formats carried inside projects
SEG-Y   LAS   DLIS   TIFF   CSV   PDF

# shape of the archive
organisation=Project ID -> File ID     grain=project level down to single files
geography=UKCS quadrants, blocks, fields   scale=petabyte class (commonly cited ~1.6 PB)

Read the anatomy rather than any single value. The identifier itself answers three questions - which company lodged it, which year, and what kind of project - before the metadata fields add completion date, contractor, licences and survey IDs. The three-state completeness line is the part most archives cannot offer: Unavailable means the requirement was acquired and then never lodged, a compliance gap with a name, while Not Acquired means the evidence never existed. Treating those two as the same 'missing' throws away the registry's entire point.

Which fields does the NDR field dictionary define?

Six documented fields form the core, split between addressing and accounting:

  • Project ID - the composite key whose segments encode company, year, type and sequence; the join key every other record hangs from.
  • File ID - the per-file handle beneath a project, the finest grain the archive indexes.
  • Classification Tag (C Tag) - what each file actually is inside the project inventory.
  • Completeness status - Available, Unavailable or Not Acquired; the audit layer.
  • Geometry / spatial location - the project's footprint against Quadrant, Block and Field polygons plus hexagon density bins, so spatial filtering is native rather than bolted on.
  • Project metadata fields - date completed, main data type, licences, contractor, description and survey ID(s).

Fields beyond the spine fold under additional fields on request below and are confirmed against live records when your sample is cut - including the expansion of the three type codes the public documentation leaves abbreviated, which is stated here rather than papered over.

Where does coverage run across geography, time and granularity?

  • Geography: the United Kingdom Continental Shelf, and nothing but it - which is the point. Every project resolves against UK offshore Quadrant, Block and Field polygons, with hexagon density bins summarising where the archive runs deepest. A jurisdiction-wide ledger under one indexing scheme beats a dozen national fragments.
  • Temporal: a continuous statutory archive rather than a series of versioned cuts. Records span the depth of the UK offshore drilling era, each carrying its own completion dates in project metadata, and the shelf keeps widening - carbon dioxide appraisal and storage licence records now arrive beside the petroleum record.
  • Granularity: project-level entries for wells, seismic surveys, hazard surveys and reports, each resolving to individual File IDs with classification tags and status. One schema serves the portfolio question and the single-well question alike.

Set against the wider catalog - 3,932 datasets averaging 7.81 - this record scores 8/10, carried by verified field documentation and statutory provenance rather than dimension count.

How is the data delivered through Datadory?

API, files, or your warehouse. Daily, weekly, or hourly.

Name the wells, quadrants, project types or years and the sample returns cut exactly to that scope - a single well's lodged record or a quadrant-wide sweep. Cadence gets settled only after the sample validates, and changing it later is a settings conversation, not a re-integration project.

Every delivery ships the field dictionary above unchanged, with the three completeness states arriving as a typed enumeration rather than free text, so an audit group-by never turns into a cleanup pass. Because the Project ID grammar encodes company, year and type, your pipeline can route and partition on the key alone before any parsing begins.

Who uses this data, and for what?

  • Decommissioning teams start from the well description projects: what was drilled, how it was completed, which contractor holds the record - and whether the evidence set reads Available or Unavailable before a P&A campaign is priced.
  • New-entry and farm-in analysts audit the lodged record for target assets, using the three-state registry as a diligence checklist rather than a vendor's assurance.
  • Geoscientists and reprocessing shops reach for legacy seismic survey projects in SEG-Y and give vintage acquisition a second life on modern processing.
  • CCS developers screen carbon dioxide appraisal and storage licence records as they join the archive, getting a statutory baseline for storage-site work.
  • ML engineers assemble log-response and seismic-characterisation corpora from LAS, DLIS and SEG-Y material that arrives keyed to real well identities.
  • Regulatory researchers and journalists cite the record itself - a statutory archive outranks any aggregation when the question is what officially happened in a basin.

Get a sample of this dataset scoped to whichever of those jobs is yours.

Which personas get the most value?

Journalists, academics and students lead the ranking at 3 of 3 - a citable statutory source with provenance no aggregator can match. Market researchers and consultants match them at 3 of 3, sizing and screening UKCS activity from official licence and well records. Data scientists and ML engineers sit at 2 of 3 with SEG-Y and LAS training material keyed to real wells, and developers and builders at 2 of 3 for metadata products built on the Project ID grammar. Investors and quant researchers close at 1 of 3 - diligence material rather than a trading feed. Persona-by-persona detail lives on the Oil & Gas Equipment & Services hub.

How does the NDR sit beside the other UK and global datasets?

The NSTA family splits by job. This archive holds the documents and measurements - well descriptions, seismic volumes, reports - lodged under statutory duty. The NSTA structured datasets registry holds the current-state geography - 922 layers tracing wellbore origins, blocks, subareas and field outlines. Use the registry to map and screen; use this archive for document-level depth on a specific well.

Against the globals, the contrast is depth versus breadth. OGIM v2.5.1 maps roughly 6.7 million infrastructure features worldwide but carries no statutory UK survey archive at all, and OPD counts platforms from orbit without reading a single well file. When the question is what did that well actually log, this is the only stop in the catalog.

What should I know before requesting a sample?

Three things, stated upfront.

First, provenance of this page: the six field definitions come from the publisher's own documentation, captured during the August 2026 research pass. Live-row examples ride in with your sample rather than being asserted here, and the expansion of the mhaz, rems and inpt type codes is confirmed against records at sampling rather than guessed.

Second, scope by named cut. A petabyte-class archive rewards precision: name the wells, quadrants, project types or years and the sample arrives shaped to the decision you are feeding, not as an undifferentiated continent of files.

Third, pair it deliberately. The NSTA structured datasets registry supplies current-state geography to screen against, OGIM supplies the global surround, and this archive supplies the evidence beneath. Start at Get a sample of this dataset; the daily, weekly, or hourly decision comes after the sample validates.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - NSTA National Data Repository (NDR), documented core
fieldtypedefinitionexample
Project IDstringUnique identifier grouping all related data for a well, survey or report set, following the fixed pattern CCCCYYYYtypeNNNN: company ID, project year, project type code and incrementing number.CCCC2020seis0001
File IDstringIdentifier for an individual file within a Project ID - the finest grain the archive indexes.-
Classification Tag (C Tag)stringClassification code describing what type of data item a given file represents within the project inventory.-
Completeness statusenumAvailability state of each information type: Available (file lodged and present), Unavailable (acquired but never lodged) or Not Acquired (never acquired). Three states, not two - the middle one is the audit signal.Available
Geometry / spatial locationgeoSpatial footprint of the project, indexed against UK offshore Quadrant, Block and Field polygons plus hexagon density bins, so projects resolve on a map rather than in a name column.-
Project metadata fieldstextDescriptive attributes on each project: date completed, main data type, licences, contractor, description and survey ID(s).-
Additional fields on request-Expansion of the mhaz, rems and inpt type codes (named against live records when your sample is cut); per-project file inventories with C Tags resolved per file; completeness rollups by operator, quadrant, year or project type; geometry extracts as coordinates.-

What teams do with it

  • Farm-in and asset due diligence Pull the lodged record for named wells before a bid: what was drilled, when it completed, which contractor ran it, and whether the evidence set reads Available or Unavailable before money moves.
  • Seismic reprocessing and reinterpretation Legacy survey projects arrive as SEG-Y, so vintage acquisitions can be reprocessed on modern stacks instead of being re-shot - the archive is where old North Sea data goes back to work.
  • Decommissioning planning Well description projects give plugging and abandonment teams the as-lodged record per wellbore, and the completeness states flag which evidence sets will need reconstructing before a campaign starts.
  • Carbon storage site screening Carbon dioxide appraisal and storage licence records are entering the archive alongside the petroleum record, giving CCS teams a statutory starting point rather than a stitching exercise.
  • Evidence-gap auditing Three-state completeness turns 'we think the data is there' into a countable matrix - Unavailable means acquired but never lodged, which is a different problem from Not Acquired and deserves a different fix.
  • Subsurface machine-learning corpora Petabyte-class, well-organised LAS, DLIS and SEG-Y material keyed to identifiable wells and surveys - rare raw material for log-response and seismic-characterisation models.

Questions buyers ask

What does one record in nsta national data repository ndr data contain?

A Project ID following the CCCCYYYYtypeNNNN pattern - company, year, project type and sequence - plus the File IDs beneath it, each with a Classification Tag, the three-state completeness status, the project's spatial footprint against Quadrant, Block and Field polygons, and metadata covering date completed, main data type, licences, contractor, description and survey IDs.

What do the three completeness statuses mean?

Available means the file was lodged and is present. Unavailable means the requirement was acquired but never lodged - a compliance gap, not a gap in nature. Not Acquired means the evidence never existed. Collapsing the last two into one 'missing' bucket discards the distinction that makes the registry useful for diligence.

How is the archive organised?

By Project ID. Each identifier bundles everything lodged for one well, seismic survey, hazard survey or report set under a fixed pattern encoding the lodging company, the year, a type code and a sequence number, so provenance survives as text before any file is opened.

Which file formats ride inside a project?

SEG-Y for seismic trace data, LAS and DLIS for well logs, TIFF for scanned sections and summary imagery, CSV for tabular exports such as reports and completeness matrices, and PDF for documents. Formats stay consistent within a project type, which is what makes bulk handling scriptable.

Does the archive cover carbon storage as well as oil and gas?

Increasingly yes. Carbon dioxide appraisal and storage licence records are being added to the archive over time alongside the conventional offshore petroleum record, indexed under the same Project ID grammar and the same completeness states, so CCS screening inherits the identical evidence-audit workflow.

Can a sample be cut to my wells, quadrants or survey years?

Yes. Name the wells, quadrants, project types or years and Datadory returns rows shaped exactly like the field dictionary above - typed enumerations included. The sample schema is the shipped schema, samples precede any commitment, and the daily, weekly or hourly decision comes after the sample validates.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing