Biotechnology · PRIDE (EMBL-EBI)

PRIDE Proteomics Archive data

Datadory delivers pride proteomics archive data covering 40,809 submitted mass-spectrometry studies from EMBL-EBI's PRIDE repository - a founding ProteomeXchange node gaining roughly 500-900 new projects a month - each carrying raw instrument files, search-engine outputs, peak lists and result files alongside CV-annotated organism, instrument, experiment-type and quantification-method fields. Delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
Global - laboratories worldwide contribute through ProteomeXchange submission routes, with a country facet on every project so contribution geography is queryable rather than anecdotal
How far back
The current archive reaches back to PXD001357, submitted 2014-10-15, and runs through present-day submissions (latest observed 2026-08-19); pre-2012 material lives in a separate legacy resource outside this feed
How fine
Project level (study), with assay/file-level records hanging beneath each project and protein identification records exposed as a further layer

What is the PRIDE Proteomics Archive?

It is the world's leading public repository for mass-spectrometry-based proteomics, delivered as queryable rows. PRIDE - PRoteomics IDEntifications Database - operates at EMBL-EBI and is a founding member of the ProteomeXchange Consortium, the arrangement under which major journals require proteomics data deposition. That mandate is why the corpus grows the way it does: 40,809 submitted projects as of August 2026, roughly 500-900 new submissions a month, and monthly submitted-data volumes of 50-100 TB.

Each project is a study, keyed by a PXD-prefixed accession and stamped with a DOI. It carries a title, a project description, free-text sample-processing and data-processing protocols, and controlled-vocabulary annotations covering organisms, organism parts, diseases, instruments, quantification methods, experiment types, software, countries and submission type. Beneath it sit the file records: raw instrument output, search-engine results, peak lists and processed result files, each classified by category with sizes and checksums. Read as browser tabs this dies in folders; resolved into rows keyed by accession, which shotgun-proteomics studies used a Q Exactive with TMT labeling on human tissue becomes a filter, not a literature review. Get a sample of this dataset and inspect real records before committing pipeline time.

What does a sample of PRIDE proteomics archive data look like?

Three real records captured during the August 2026 pass - the archive's oldest current project, its newest observed submission, and one file record from beneath a project:

# Row 1 — the oldest project in the current archive, complete record
accession           : PXD001357
title               : Direct evidence of milk consumption from ancient human
                      dental calculus, St Helena
doi                 : 10.6019/PXD001357
submissionDate      : 2014-10-15
submissionType      : COMPLETE
organisms           : Homo sapiens (human)
instruments         : LTQ Orbitrap Elite | Q Exactive
experimentTypes     : Shotgun proteomics
projectTags         : Technical | Metaproteomics
totalFileDownloads  : 29298

# Row 2 — the newest submission observed in the August 2026 research pass
accession           : PXD082855
title               : Mapping the GW182 CIM1 binding site on CNOT1(800-999)
                      by HDX-MS
submissionDate      : 2026-08-19

# Row 3 — a file record hanging beneath a project
fileCategory        : Search engine output file URI
fileName            : F116783_JH8_QE.dat
projectAccession    : PXD001357

Read what the spread proves. Twelve years separate those two projects, yet the record shape never moved: accession, DOI-bearing title, dated submission, CV-annotated science. The first row shows why the annotations earn their keep - two instruments named precisely enough to compare, an experiment type drawn from a controlled vocabulary, and a download counter (29,298) that no citation metric captures. The third row shows the grain beneath the grain: a project is not one blob but a classified set of files, and fileCategory is what tells a spectral-library builder from a reanalysis team which layer to take. Example rows are illustrative of the record shape; field-level detail follows below.

What fields does the PRIDE Proteomics Archive include?

Nineteen core fields cover project and file records, every definition checked against live records during the August 2026 research pass. They do three jobs.

Identity and dating: accession anchors the row on the PXD identifier every downstream resource cites, doi gives it a permanent citable handle, and submissionDate plus publicationDate place it on the timeline. submissionType grades completeness - COMPLETE, PARTIAL or PRIDE - which is a ready-made filter for whether a deposit holds everything a reanalysis needs.

Scientific context: the controlled-vocabulary block that makes this corpus joinable - organisms from NEWT taxonomy, instruments from the MS ontology, experimentTypes and quantificationMethods from PRIDE's own term lists, plus submitter keywords and projectTags. Because these arrive as vocabulary terms rather than free text, an instrument comparison across decades is a GROUP BY instead of a string-matching archaeology dig.

Methods and weight: the two protocol texts preserve how samples were prepared and how spectra were processed - the reproducibility record papers summarize away - while fileCategory, fileName, fileSizeBytes and totalFileDownloads describe and size the file estate beneath each project. Entries flagged in the footnote fold under additional fields on request.

Which fields arrive only on request?

Beyond the nineteen-column core, four extensions extend the same universe sideways; each gets mapped to your use case at sampling:

  • Checksums and public location records on files - present on file rows during verification; they ship when your pipeline needs integrity checking or mirroring.
  • Organism-part and disease annotations - further CV-param blocks carried on subsets of projects, folded in when clinical slicing matters.
  • Protein identification records - the protein-level layer beneath projects, joined on by accession when identifications rather than studies are the unit of work.
  • Reanalysis links and SDRF metadata exports - available where submitters provided them; SDRF availability in particular varies project by project, so it is confirmed per slice rather than promised blind.

Name the ones your models need when you request a sample and the extract comes back cut to exactly that shape.

What geography, time range, and granularity does the dataset cover?

Geography: global by construction. Laboratories worldwide deposit through ProteomeXchange routes, and a country facet rides on every project - so contribution geography is a queryable column rather than an inference from author affiliations. Organism, more than country, is usually the axis that matters: coverage tracks what the world's mass spectrometers have been pointed at.

Temporal: the current archive runs from PXD001357, submitted 2014-10-15, forward through present-day submissions - the newest observed during the August 2026 pass landed on 2026-08-19. One structural honesty note: material older than the ProteomeXchange era sits in a separate legacy resource outside this feed, so decade-spanning backtests should say so upfront. Freshness is not a constraint on your side either way - deliveries land daily, weekly, or hourly regardless of the underlying cadence.

Granularity: the project is the unit - a study with its protocols, annotations and DOI - and two finer layers hang beneath it: file-level records classified by category and sized in bytes, and protein identification records exposed as a further stratum. Fix the grain before scoping a request, because mixing studies with their constituent files inflates counts fast.

Scale check against Datadory's wider catalog - 1,744 datasets across 159 viable industries, average quality score 7.81 - this slice scores 8/10 with field definitions verified during research, and its growth rate means any snapshot understates next quarter's holdings.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

You pick the channel and the cadence; the field dictionary above travels unchanged through all three. Rows arrive flattened - one project per row, file records normalized beneath, CV terms kept as clean strings - so joining against your own target list, instrument inventory or disease panel is a join statement rather than a parsing project. Cadence changes are a settings conversation, not a re-integration, and a sample cut to your named organisms, instruments or date range comes first either way.

Who uses this data, and for what?

  • Biomarker discovery panels - filter by organisms, disease annotation and quantificationMethods to assemble candidate cohorts before reading a single paper; the TMT-versus-label-free split alone collapses thousands of studies into the comparable subset.
  • Instrument fleet analytics - the MS-ontology instruments block turns 40,809 projects into a market-share history of mass spectrometers: which platforms ran which experiment types, when, at what volume. Vendors and procurement teams price capacity decisions off that curve.
  • Method reproducibility audits - sampleProcessingProtocol and dataProcessingProtocol text plus the submissionType completeness grade let reviewers check whether a paper's stated workflow matches what was actually deposited. Discrepancy becomes a query, not a peer-review hunch.
  • Spectral library building - fileCategory separates raw instrument output from search-engine results and peak lists, so library builders pull the right stratum without downloading archives blind.
  • Metaproteomics and ancient-protein research - projectTags isolates technical and metaproteomics deposits (dental-calculus milk proteins included) that generalist catalogs bury.
  • Research-trend intelligence - submissionDate joined with experimentTypes yields a monthly-resolution series on what the field is actually measuring; see competitive intel workflows.

Which personas get the most value?

Ranked by relevance in Datadory's persona tagging:

  1. Bioinformaticians & computational biologists - accession-keyed project and file records that join outward on DOIs and NEWT taxonomy without fuzzy matching, so reproducing a paper's inputs is a filter instead of a correspondence chain.
  2. Data scientists & ML engineers - tens of thousands of labeled projects (organism, instrument, method, tag) as supervision for retrieval and cohort-matching models; the data scientists working in biotechnology briefing goes deeper.
  3. Competitive intelligence & product teams - instruments and software annotations read which platforms dominate real depositions when judging vendor positioning; see developers building in biotechnology for the product angle.
  4. Journalists, academics & students - DOI-carrying records citable at source, with totalFileDownloads as a ready-made impact story no citation count tells.
  5. Investors & diligence teams - submission-volume trends by method and instrument as a demand signal on the proteomics toolchain before the term sheet moves.

Which datasets sit next to this one?

The Biotechnology shelf splits the proteomics job by vantage point, and this record owns the deposition-evidence layer. Human Protein Atlas says where a protein appears in the body; the archive here supplies the mass-spec evidence that a peptide signature was actually measured - pair them and a tissue-restriction claim gains its instrument-grade receipt. UniProt provides the protein identities every identification resolves against, and RCSB Protein Data Bank supplies the experimental structures a validated interaction must fit into. European Nucleotide Archive covers the genomics side of the same mandate-and-deposit world, with a matching accession-first record grammar. GWAS Catalog adds the trait associations worth checking when a proteomic hit needs genetic corroboration. All of them, ranked and cross-linked, sit in the best biotechnology datasets shortlist; the pooled biotechnology data hub holds the full view, and the PRIDE (EMBL-EBI) source profile covers the repository itself.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - nineteen verified fields on PRIDE project and file records; extensions fold under additional fields on request
fieldtypedefinitionexample
accessionstringUnique ProteomeXchange/PRIDE project identifier with a PXD prefix - the primary key every downstream resource cites.PXD001357
titlestringTitle of the submitted proteomics study as the submitters wrote it.Direct evidence of milk consumption from ancient human dental calculus, St Helena
projectDescriptiontextFree-text description of the study's aims and design.
sampleProcessingProtocoltextFree-text description of how samples were prepared before MS analysis - the reproducibility half of the methods section.
dataProcessingProtocoltextFree-text description of spectra processing, database search and statistical analysis.
doistringDigital object identifier minted for the dataset - the citable handle that outlives URL churn.10.6019/PXD001357
submissionDatedateDate the project was submitted to PRIDE.2014-10-15
publicationDatedateDate the project was made publicly visible.
submissionTypeenumProteomeXchange submission completeness class: COMPLETE, PARTIAL or PRIDE - a ready-made flag for whether the deposit holds everything.COMPLETE
organismstextControlled-vocabulary list of source organisms from NEWT taxonomy, e.g. Homo sapiens (human).Homo sapiens (human)
instrumentstextControlled-vocabulary list of mass spectrometers used, from the MS ontology - Q Exactive, LTQ Orbitrap Elite and their kin.Q Exactive
experimentTypestextControlled-vocabulary experiment type, e.g. Shotgun proteomics (PRIDE:0000429).Shotgun proteomics
quantificationMethodstextControlled-vocabulary quantification approach, e.g. TMT or Label free.
keywordstextSubmitter-supplied keywords describing the study.Human | Beta-lactoglobulin | Dental calculus
projectTagstextCategorization tags such as Biological, Biomedical or Technical.Technical | Metaproteomics
fileCategorytextControlled-vocabulary class of an associated file - raw instrument file, search engine output, peak list or result file.Search engine output file URI
fileNamestringName of the associated data file within the project.F116783_JH8_QE.dat
fileSizeBytesintegerSize of the file in bytes - the column that lets you budget storage before pulling anything.
totalFileDownloadsintegerCumulative download count for the project - a usage signal no citation count gives you.29298

What teams do with it

  • Biomarker discovery panels Filter projects by organism, disease annotation and quantification method to assemble candidate cohorts before reading a single paper - the TMT-versus-label-free split alone narrows thousands of studies to the comparable subset.
  • Instrument fleet analytics The MS-ontology instrument block turns 40,809 projects into a market-share history of mass spectrometers: which platforms ran which experiment types, when, at what volume.
  • Method reproducibility audits Sample-processing and data-processing protocol text plus the submissionType completeness class let reviewers check whether a paper's stated workflow matches what was actually deposited.
  • Spectral library building File-category classification separates raw instrument output from search-engine results and peak lists, so library builders can pull the right stratum without opening archives blind.
  • Metaproteomics and ancient-protein research Project tags isolate technical and metaproteomics deposits - dental-calculus milk proteins included - that general genomics catalogs bury.
  • Research-trend intelligence Submission dates plus experiment types give equipment vendors, CROs and investors a monthly-resolution series on what the field is actually measuring, not what it says it measures.

Questions buyers ask

How big is the PRIDE Proteomics Archive?

40,809 submitted projects as of August 2026, growing by roughly 500-900 new submissions a month at monthly submitted-data volumes of 50-100 TB, with tens of millions of individual files beneath the project layer.

How far back do the records go?

The current archive starts at PXD001357, submitted 2014-10-15, and runs to present-day submissions - the newest observed in the August 2026 pass landed 2026-08-19. Pre-ProteomeXchange material sits in a separate legacy resource outside this feed, so decade-spanning backtests should be scoped explicitly.

What fields does each record carry?

Nineteen core fields across project and file rows: PXD accession, title, description, both processing-protocol texts, DOI, submission and publication dates, submission type, and controlled-vocabulary blocks for organisms, instruments, experiment types, quantification methods, keywords, tags and file categories, plus file size and cumulative download counts.

Can I filter by instrument and quantification method?

Yes, and that is the point of the vocabulary blocks. Instruments come from the MS ontology (Q Exactive and LTQ Orbitrap Elite are distinct terms), quantification methods distinguish TMT from label-free, and experiment types such as Shotgun proteomics carry PRIDE term IDs - so platform and method comparisons resolve as filters rather than string matching.

Does the dataset include the actual spectra or just metadata?

Both layers exist. Project records carry the study metadata and annotations, while file records beneath each project classify raw instrument files, search-engine outputs, peak lists and result files with names, categories and byte sizes - the dictionary distinguishes them so you can scope exactly which stratum a delivery carries.

Which organisms does the archive cover?

Anything submitted: organism annotations use NEWT taxonomy on every project, from Homo sapiens through metaproteomics communities, with a country facet riding alongside so contribution geography is queryable. Coverage tracks what the world's mass spectrometers have been pointed at rather than any jurisdiction.

How often can deliveries be scheduled?

Daily, weekly, or hourly - your call, changeable later without re-integration. Monthly suits trend work; weekly keeps cohort lists current through the 500-900 monthly new submissions; hourly serves monitoring pipelines watching for deposits on specific organisms or instruments.

Who uses pride proteomics archive data?

Bioinformaticians reproducing published analyses from accession-keyed records; data scientists training retrieval and cohort-matching models; competitive-intel teams reading instrument adoption from real depositions; journalists citing DOI-stamped studies; and diligence teams tracking submission trends across the proteomics toolchain.

Notes on this record

  • Completeness is a field, not a guess The submissionType grade (COMPLETE / PARTIAL / PRIDE) tells you whether a deposit holds everything before you plan a reanalysis around it.
  • Download counts beat citation lag PXD001357 shows 29,298 cumulative downloads - a usage signal on day one that citation metrics need years to produce.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing