Oil & Gas Equipment & Services · Kaggle (Yohanes Nuwara)
Multimodal RAG for Oil and Gas (Kaggle Dataset)
Datadory delivers oil & gas equipment services data covering multimodal rag for oil and gas kaggle dataset data: eight genuine petroleum PDFs - a 94-page Sleipner completion report, its drilling programme, discovery and SCAL studies, a Duvernay PVT report, an SPE paper and the Drilling Data Handbook - packaged for vision-language retrieval. Delivered daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Two basins, deliberately chosen: the Sleipner field in the Norwegian North Sea (wells 15/9-15 and 15/9-19 A, plus the discovery wells behind Discovery_Report.pdf) and the Duvernay field in Canada. Depth over breadth - a corpus you can know completely rather than sample thinly.
- How far back
- Paperwork from the late 2000s - the Sleipner final well report is dated 12 August 2009 - packaged through seven compilation versions released between 7 December 2024 and 6 January 2025. A frozen benchmark, not a feed.
- How fine
- Document-level: each entry is a complete multi-page report with embedded tables, charts and schematics, and the effective retrieval grain is the page image with layout intact.
What is the Multimodal RAG for Oil and Gas dataset?
It is a 337 MB argument that petroleum engineering needs its own retrieval stack. Multimodal RAG for Oil and Gas packages eight real documents - operator reports and reference texts totalling 337,264,957 bytes - assembled by petroleum data scientist Yohanes Nuwara for building retrieval-augmented generation pipelines that see pages rather than strings.
The core is the Sleipner field in the Norwegian North Sea. Four of the eight documents concern it: the final well completion report for well 15/9-15, a 94-page Statoil document dated 12 August 2009 that walks through wellbore sequences, formation evaluation, plugbacks, contractors and cost/time distribution; the drilling programme for the same well; a discovery report compiled from several Sleipner discovery wells; and a compressed SCAL special core analysis report for well 15/9-19 A. A PVT report from the Duvernay field in Canada, an SPE paper by Luo (1994), the Drilling Data Handbook and the Norwegian-language volume Olje og Gass Fra Felt til Raffineri round out the set.
What makes the collection pointed is the publisher's stated recipe: index the documents with ColPali, retrieve whole pages for a query, then generate the answer from those pages with Qwen2-VL-2B or Llama Vision. Seven compilation versions shipped between 7 December 2024 and 6 January 2025, with two community notebooks written against them. Inside Datadory's 1,744-dataset catalog this record anchors the document-intelligence end of the fourteen-dataset Oil & Gas Equipment & Services set. Amaze yourself with what is in it, then get the pages that matter:
What do sample rows look like?
One document, three grains - the corpus in miniature:
# the retrieval unit -- one page, not one string
corpus_size : 337,264,957 bytes (~337 MB)
documents : 8 PDFs
retrieval_grain : page image + layout, queried directly
# FINAL WELL REPORT (94 pages), Drilling Licence PL046BS
wells : NO 15/9-F-15 A/B/C
dated : 12.08.2009
sections : general well data; HES&Q incidents by service
and company; time distribution; cost;
formation evaluation; coring summary;
pressure points; geological formation tops
# Appendix 4: Contractors list (Sleipner well 15/9)
cementing : Halliburton
rig_operations : Maersk Contractors
liner_hanger_equipment : Weatherford Norge AS
mud_logging : Schlumberger
mwd : Schlumberger
electric_wireline_logging: Schlumberger
drilling_fluids : Halliburton
directional_drilling : Schlumberger
subsea_wh_xmas_tree : Vetco Gray
rov_systems : Oceaneering
directional_survey : Gyrodata
# mechanical plug record, wellbore NO 15/9-F-15 C
date : 28 Jan 2009 size : 10 3/4 in plug_type : RTTS
set_depth : 319.0 mMD tagged_depth : 210.0 mMDRead what those rows prove. A retrieval system asked who cements wells on Sleipner? finds Halliburton in Appendix 4; asked who ran downhole measurement it finds Schlumberger holding mud logging, MWD and wireline at once; asked what was set where, it finds a 10 3/4 inch RTTS plug at 319.0 mMD with a tagged depth of 210.0 mTD on 28 January 2009. These are structured facts living inside unstructured PDFs - exactly the gap multimodal RAG claims to close, and exactly why practising on real operator paperwork beats practising on scanned menus.
Sample rows above show document-level structure observed during the August 2026 research pass; page-level tables ship inside your sample.
Which fields does the dataset dictionary define?
Six entries carry the release. COMPLETION_REPORT.pdf is the heavyweight: the final well report for well 15/9-15 on drilling licence PL046BS, ninety-four pages running from general well data through HES&Q incidents, time distribution, cost, formation evaluation, coring summary, pressure points and geological formation tops. DRILLING_PROGRAMME_1.pdf holds the plan for the same well, which turns the pair into a plan-versus-as-executed comparison waiting to be automated.
Discovery_Report.pdf widens from one wellbore to field scale, compiling several Sleipner discovery wells into one narrative. SCAL_15-9-19_A_0035_compressed.pdf goes the other direction - down to special core analysis for well 15/9-19 A, the laboratory grain reservoir models are calibrated against. WF-Gas-Condensate-REC-Sample (CL-63169).pdf crosses the Atlantic for Duvernay gas-condensate PVT behaviour in Canada.
The last entry is the reference shelf: an SPE paper by Luo (1994), the Drilling Data Handbook and the Norwegian volume Olje og Gass Fra Felt til Raffineri, giving a retrieval system domain vocabulary to draw on alongside operator paperwork.
Definitions above are verified against the published record during the August 2026 research pass. Anything adjacent folds under additional fields on request rather than being guessed here - per-document page counts, the two community notebooks, the deeper file inventory and the seven-version changelog all confirm when your sample is cut.
Where does coverage run across geography, time and granularity?
Geography: two basins, deliberately chosen. The Sleipner field in the Norwegian North Sea supplies wells 15/9-15 and 15/9-19 A plus the discovery wells behind the discovery report; the Duvernay in Canada contributes one PVT study. This is a corpus you can know completely, not one you sample.
Temporal: the paperwork dates from the late 2000s, anchored by the final well report's 12 August 2009 date, while the packaging itself ran through seven versions between 7 December 2024 and 6 January 2025. Nothing has moved since. Treat it as a frozen benchmark, and pair it with a moving feed when a pipeline needs fresh input.
Granularity: document-level, with each PDF a complete multi-page report carrying embedded tables, charts and schematics - and the effective retrieval grain being the page image with its layout intact. Set against the wider catalog - 1,744 datasets averaging 7.81 - this record scores 5/10, carried by verified provenance and genuine operational texture, held back by having no tabular layer at all: every fact lives inside a page, and extracting it is your model's job. That difficulty is the point of the benchmark.
How is the data delivered through Datadory?
API, files, or your warehouse. Daily, weekly, or hourly.
Because the corpus is static, cadence matters less than plumbing: pick the channel your retrieval stack already speaks and the eight documents arrive intact - PDFs preserved as PDFs, nothing pre-chunked into someone else's idea of a paragraph. Chunking strategy is where retrieval accuracy lives or dies, so we leave it to you; ask and the sample arrives pre-split at page boundaries to match a ColPali-style indexer.
Every delivery ships the field dictionary above plus sample rows for validation. Name the documents you care about - Sleipner only, or the full eight - and the sample comes back shaped to them before any commitment.
Who uses this data, and for what?
- Vision-language retrieval tuned to petroleum documents - index pages with ColPali, answer with Qwen2-VL-2B or Llama Vision, and measure how far a general-purpose retriever drifts on drilling vocabulary, formation names and plug records.
- Service-company extraction - eleven named contractor assignments on one well, HES&Q incidents broken down by service and company, cost and time distributions: supply-chain and market-sizing evidence already sitting inside real reports.
- Petrophysics and PVT assistants - a SCAL special core analysis report and a Duvernay gas-condensate PVT study give question-answering systems laboratory-grade material to be graded against.
- Plan-versus-as-executed copilots - the drilling programme and the completion report cover the same well, so a copilot answers what changed without leaving the corpus.
- Teaching and case-study material - dated operator reports with named contractors and tagged plug depths turn well-construction instruction into something students can interrogate.
Each job maps onto a persona page: data scientists, developers builders, and investors quants reading document-AI tooling adoption as demand signal.
Which personas get the most value?
Data scientists and ML engineers lead, because the binding constraint on document AI in energy is not architecture - it is evaluation data with real operational texture, and eight documents you can fully verify beat ten thousand you cannot.
Developers and builders get a small, self-contained corpus: about 337 MB, PDF-native, no join keys to reconstruct, runnable end to end on a workstation. The interesting engineering decision - page-level versus section-level chunking - stays yours.
Market researchers and consultants read the contractor lists as a supply-chain census of one North Sea well: Halliburton on cementing and drilling fluids, Schlumberger on the logging string, Maersk Contractors on the rig, Oceaneering on ROV support.
Academics and students get citable primary sources rather than summaries - a dated Statoil well report and a peer-reviewed SPE paper. Persona-by-persona detail lives on the industry hub at /industries/oil-gas-equipment-services.
How does it sit beside the rest of the oil & gas equipment & services catalog?
This record owns the unstructured corner of the industry's fourteen-dataset catalog, and its neighbours mark the boundary precisely. NORA 3D Multimodal Oil & Gas Dataset is also multimodal, but geometrically so - annotated point clouds, CAD and P&IDs where this corpus is prose, tables and schematics. GainEnergy Oil & Gas Engineering Dataset offers synthetic problem statements with clean labels where these documents are authentic and unlabeled.
Sensor streams such as the refinery centrifugal compressor year answer how is the machine running; this corpus answers how was the well built, and who built it. Used together they cover a facility from paperwork to vibration. The wider ranking sits on best oil & gas equipment services datasets.
What should I know before requesting a sample?
Three things worth settling upfront.
First, scope honestly. Eight documents is a benchmark, not a library. If your question needs many wells, many fields or current-year reporting, say so when requesting the sample and we will scope companions from the catalog rather than overselling this record.
Second, expect to do the structuring. There is no tabular layer here - contractor lists, plug records and pressure points arrive inside pages, and pulling them out is precisely the capability you are testing. Ask for the sample pre-split at page boundaries if that matches your indexer.
Third, weigh the record's own signals: seven compilation versions between 7 December 2024 and 6 January 2025, and a 0.75 usability score on the platform's own rubric, which mostly reflects documentation gaps rather than content problems. Both argue for validating against real pages before committing - which is what the sample is for.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
COMPLETION_REPORT.pdf | document | Final well completion report for well 15/9-15, Sleipner field, Norwegian North Sea - the 94-page Statoil document carrying wellbore sequences, formation evaluation, plugbacks, contractors and cost/time distribution. | 94 pages, dated 12.08.2009 |
DRILLING_PROGRAMME_1.pdf | document | Drilling programme for well 15/9-15, Sleipner field - the plan document that pairs against the as-drilled completion report for plan-versus-executed analysis. | well 15/9-15 |
Discovery_Report.pdf | document | Discovery report compiled from several discovery wells of the Sleipner field, widening the corpus from one wellbore to field-scale context. | multiple Sleipner discovery wells |
SCAL_15-9-19_A_0035_compressed.pdf | document | Petrophysical SCAL (special core analysis) report for well 15/9-19 A, Sleipner field - the laboratory-grain rock properties behind reservoir modelling. | well 15/9-19 A |
WF-Gas-Condensate-REC-Sample (CL-63169).pdf | document | Pressure-volume-temperature (PVT) report for a Duvernay field gas condensate sample in Canada, adding fluid-behaviour evidence to the North Sea core. | Duvernay field, Canada |
SPE paper (Luo 1994) and reference books | document | An SPE paper by Luo (1994), the Drilling Data Handbook and the Norwegian-language petroleum volume Olje og Gass Fra Felt til Raffineri - the reference shelf beside the operator paperwork. | Luo (1994); Drilling Data Handbook |
Additional fields on request | - | Per-document page counts beyond the 94-page final well report, the two community notebooks written against the release, the full file inventory beneath the archive's top level, and the version-by-version changelog across the seven published compilations. | - |
What teams do with it
- Vision-language retrieval tuned to petroleum documents Index pages with ColPali, answer with Qwen2-VL-2B or Llama Vision, and measure how far a general-purpose retriever drifts on drilling vocabulary, formation names and plug records.
- Service-company extraction Eleven named contractor assignments on one Sleipner well, HES&Q incidents broken down by service and company, cost and time distributions: supply-chain evidence already sitting inside real reports.
- Petrophysics and PVT assistants A SCAL special core analysis report and a Duvernay gas-condensate PVT study give question-answering systems laboratory-grade material to be graded against.
- Plan-versus-as-executed copilots The drilling programme and the completion report cover the same well, so a copilot can answer what changed between plan and outcome without leaving the corpus.
- Teaching and case-study material Dated operator reports with named contractors and tagged plug depths turn a lecture on well construction into something students can interrogate directly.
Questions buyers ask
What does the Multimodal RAG for Oil and Gas dataset contain?
Eight petroleum documents totalling about 337 MB: a 94-page Sleipner final well completion report, the drilling programme for the same well, a Sleipner discovery report, a SCAL petrophysical report for well 15/9-19 A, a Duvernay PVT report, an SPE paper by Luo (1994), the Drilling Data Handbook and a Norwegian petroleum industry volume.
Why is it described as multimodal rather than plain text?
Because the working unit is the page, not the string. These reports carry tables, casing and completion schematics, charts and stamp blocks that survive OCR poorly, so the publisher's stated recipe indexes page images with ColPali and generates answers from retrieved pages using Qwen2-VL-2B or Llama Vision.
Which wells and basins do the documents cover?
The Sleipner field in the Norwegian North Sea anchors the set - wells 15/9-15 and 15/9-19 A plus several discovery wells - and the Duvernay field in Canada contributes one gas-condensate PVT report. Two basins and four documented wells, chosen for report quality rather than geographic spread.
Is the corpus a time series or a fixed snapshot?
Fixed. The paperwork dates from the late 2000s - the Sleipner final well report is dated 12 August 2009 - and the compilation itself ran through seven versions between 7 December 2024 and 6 January 2025. Nothing moves; pair it with a live feed if your pipeline needs movement.
How large is the collection?
337,264,957 bytes - about 337 MB - across eight PDFs. That is small enough to index in full and evaluate end to end on a single workstation, which is the design intent: a corpus you can know completely while tuning retrieval and answering strategies.
Who built the dataset?
Petroleum data scientist Yohanes Nuwara assembled and maintained the seven-version compilation, publishing it with a stated ColPali-plus-vision-language-model workflow and two accompanying community notebooks. Its users skew toward ML teams adapting retrieval-augmented generation to energy documents and consultants reading contractor line items as supply-chain evidence.
Notes on this record
- Pages, not paragraphs The publisher's stated workflow retrieves whole report pages with ColPali and answers from them with Qwen2-VL-2B or Llama Vision - so tables and schematics survive into the answer instead of being flattened by OCR.
- Operational texture, on the record Appendix 4 of the Sleipner final well report names eleven contractor assignments - Halliburton cementing, Schlumberger across the logging string, Maersk Contractors on the rig, Gyrodata on surveys. Generic corpora never carry this.
- Two basins, one winter Sleipner and Duvernay paperwork packaged through seven versions between 7 December 2024 and 6 January 2025, then frozen. Treat it as a benchmark; borrow a live feed when a pipeline needs movement.
- Small enough to finish About 337 MB across eight PDFs: the entire corpus fits one retrieval experiment, end to end, on a workstation - no cluster required to evaluate a change.
- Sample policy Samples ship cut to the documents and page ranges you name, with the field dictionary above unchanged plus sample rows for validation.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.