UniProt REST API - Protein Sequence & Function
Datadory delivers UniProt REST API protein sequence and function data covering 149,810,139 knowledgebase entries - 575,503 of them manually reviewed - each carrying function text, subcellular location, domain coordinates, GO terms, EC numbers and 300-plus cross-references, flanked by cluster, archive and proteome layers. Delivered as API pulls, files or warehouse writes, daily, weekly or hourly.
What is UniProt REST API - Protein Sequence & Function?
This is the reference protein knowledgebase of molecular biology, delivered as one addressed dataset inside Datadory's life sciences tools & services slice. Current vintage 2026_02 (10 June 2026): 149,810,139 UniProtKB entries, of which 575,503 carry hand-curated Swiss-Prot verdicts and the remaining 149,234,636 ride computationally annotated TrEMBL rails - flanked by 381,103,551 UniRef clusters, 1,158,429,795 archived non-redundant UniParc sequences and 1,086,758 organism proteomes. Four corpora, one accession spine.
Eighteen verified fields define every row: recommended and alternative names, gene symbols, organism and taxon ID, sequence length and computed mass, evidence-tagged function and subcellular-location text, per-feature coordinates for domains, signal peptides, disulfide bonds and modified residues, GO terms split by ontology branch, EC numbers, and cross-references out to structures, chemistry, drugs and pathways - with audit dates, a 0-to-5 annotation-completeness score and the five-level protein-existence tier closing each record.
It scores 10 out of 10 on Datadory's rubric against a catalog averaging 7.81 across 1,744 datasets - a mark only 145 records earn. Within the slice it plays the protein half beside PubChem PUG REST API's compound half. Get a sample of this dataset cut to your own target list.
What do sample rows from the dataset look like?
One reviewed entry per row, identity and curation history pinned side by side. Rabbit mitochondrial aspartate aminotransferase, straight from the verification capture:
# delivered grain: one protein entry per row
primaryAccession : P12345
uniProtkbId : AATM_RABIT
entryType : UniProtKB reviewed (Swiss-Prot)
protein_name : Aspartate aminotransferase, mitochondrial
gene : GOT2
organism : Oryctolagus cuniculus taxon_id : 9986
seq_length : 430 aa
first_public : 1989-10-01
last_annotation : 2026-06-10 version : 149
# two computed-layer rows from the same delivery, extremes of the size range
A0A0C5B5G6 MOTSC_HUMAN Mitochondrial-derived peptide MOTS-c Homo sapiens 16 aa 2,175 Da
A0A1B0GTW7 CIROP_HUMAN Ciliated left-right organizer metallopeptidase (EC 3.4.24.-) Homo sapiens 788 aa 85,397 Da
# identifier translation layer
P12345 -> UniRef100_P12345 # accession-to-cluster mapping, 100+ identifier types supportedRead the top row as three decisions already made for you. The evidence tier travels as a column, so curated and computationally annotated never blend mid-query. The curation history arrives as two dates - public in October 1989, still being worked in June 2026, thirty-seven years of continuous curation compressed into version 149. And geometry ships as data: length and mass sit as integers ready to filter on, not prose waiting to be parsed.
The two smaller rows prove the schema holds at the extremes - a 16-amino-acid peptide and a 788-residue metallopeptidase land in the same columns as the enzyme above, their EC class riding the name. The last row is the translation layer doing its job: one accession in, its cluster home out, one row among mappings spanning 100-plus external identifier types.
Every value above comes from the catalog's August 2026 verification pass, not a mock-up.
Which fields does the dataset include?
Eighteen documented fields define the entry, each definition verified against live output during the August 2026 research pass - a bar only 145 of 1,744 cataloged datasets clear outright. They split into six jobs:
- Identity and nomenclature:
primaryAccession,uniProtkbId,entryTypeand the recommended full name pin one stable address per protein, with gene symbols riding alongside. - Organism:
organism.scientificNameplus the NCBI taxon ID make every species-and-region cut a join rather than a guess. - Sequence geometry:
sequence.lengthand the computedsequence.mass, shipped as integers. - Annotation: function and subcellular-location comment blocks carrying evidence codes, per-feature coordinates (
ft_domain,ft_signal,ft_disulfid,ft_mod_res), Gene Ontology terms split by branch and Enzyme Commission numbers. - Cross-references:
xref_pdb,xref_alphafolddb,xref_chembl,xref_drugbank,xref_reactomeand 300-plus pointer-field siblings. - Provenance and quality: first-publication and last-annotation audit dates, the 0-to-5
annotation_scoreand theprotein_existencetier.
The full dictionary follows in tabular form below; anything family-specific folds into your sample on request.
What does coverage look like across geography, time and granularity?
- Geography - global by construction: sequences from all domains of life across 1,086,758 proteomes, contributed and curated internationally. No record carries a jurisdiction flag; species-and-region slices come from joining the taxonomy columns outward.
- Temporal - entries integrated since 1986, with the oldest flat-file date in the verification capture reading 1 October 1989. Every row indexes its own first-publication and last-annotation dates, and each delivery carries its vintage stamp (2026_02, June 2026), so the state of the corpus at a chosen date reconstructs cleanly.
- Granularity - three levels deep: whole-protein entries, then per-feature coordinates hanging off each sequence, then individual cross-reference rows beneath those. Wide scans (every reviewed human kinase) and narrow ones (every disulfide bond annotated since a given date) resolve without touching non-matching rows.
Set against the wider catalog - 1,744 datasets averaging 7.81 - the 10/10 prices in verified definitions, evidence tiers kept separable, and four companion corpora behind one accession spine. Where it ranks: best life sciences tools services datasets.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
You pick the channel and cadence; parsing the annotation nesting, keeping reviewed apart from computed and flattening features into joinable columns stay our problem. Rows arrive accession-keyed, so this month's delivery appends cleanly onto last month's and joins whatever target, compound or expression tables you already hold.
Hourly suits identifier-resolution services, where a stale accession map breaks lookups downstream. Daily fits screening pipelines triaging newly annotated proteins overnight. Weekly matches the rhythm most protein-annotation workflows actually run at. Whichever you choose, the eighteen-field dictionary travels unchanged. A sample ships first either way - real rows for your organisms, families and annotation floors before any commitment.
Who uses this data, and for what?
Six jobs the protein reference settles outright:
- Target triage and prioritization - pull every reviewed member of a family filtered by GO branch, EC number and domain architecture to see which targets carry known active sites and which are still structurally dark.
- Function-prediction model training - reviewed entries supply curator-approved labels while
annotation_scoreandprotein_existencegrade label confidence; see ml model training. - Antibody and assay design - signal peptides, transmembrane regions, extracellular domains and disulfide bonds arrive as coordinates, so epitope choices happen before anyone orders a reagent.
- Identifier resolution - accession-to-cluster mappings and 100-plus external identifier types resolved into one keyed lookup; see api integration.
- Knowledge-graph construction - protein nodes pre-linked to structures, drugs and pathways on identifiers both sides already share, ready to extend outward rather than reconciled by hand.
- Citation-grade verification - the protein a paper names resolves to accession, evidence tier and annotation dates; see citation grade research.
Each job maps to a persona below, and the sample validates whichever one you came for.
Which personas get the most value?
Data Scientists & ML Engineers lead fit: decades of curator-labeled proteins with confidence dials attached is the combination protein-level models rarely get in one place - see life sciences tools services data for data scientists. Developers & Data-Product Builders wire stable accessions into bio apps, because an accession survives redesigns far better than a name ever will - see life sciences tools services data for developers builders. Competitive Intelligence & Product Teams size protein families and mechanism crowding before committing to a position. Market Researchers & Consultants cite the reviewed-entry counts when sizing proteomics research markets. Journalists, Academics & Students get citable answers with the evidence attached to each claim.
Persona fit has edges worth naming: this annotates proteins; it does not measure expression, activity or patient outcomes. Those live in neighboring catalogs - tissue context at Human Protein Atlas, transcriptomic evidence at NCBI GEO, measured activities at ChEMBL. Pair accordingly rather than expecting one feed to cover the whole bench-to-bedside chain.
Which cards pair with this dataset?
Cards worth opening next:
- Life Sciences Tools & Services data hub - the pooled industry view this record files under, three primaries deep.
- Best life sciences tools services datasets - where this perfect score ranks inside its industry.
- UniProt source profile - consortium background and the nine named deliverables cut from this corpus.
- PubChem PUG REST API - Chemical Structures & Properties - the compound half of the slice, quality 10 beside quality 10.
- PubChem PUG REST API vs UniProt REST API - small-molecule identity against protein annotation, scored honestly.
- UniProt vs CZ CELLxGENE Discover - protein-level annotation against single-cell expression.
- UniProt under Biotechnology - the same corpus filed with its pooled-industry neighbors.
Then get a sample of this dataset - real rows cut to your organisms, families and annotation floors before any commitment.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
primaryAccession | string | Primary accession anchoring the entry; the stable join key every delivery keys on. | P12345 |
uniProtkbId | string | Entry mnemonic identifier within its organism. | AATM_RABIT |
entryType | enum | Reviewed status separating manually curated Swiss-Prot records from computationally annotated TrEMBL ones. | UniProtKB reviewed (Swiss-Prot) |
proteinDescription.recommendedName.fullName.value | string | Recommended protein name, with alternative names riding beside it. | Aspartate aminotransferase, mitochondrial |
genes[].geneName.value | string | Gene symbol(s) associated with the protein, plus synonyms across naming conventions. | GOT2 |
organism.scientificName | string | Source organism scientific name. | Oryctolagus cuniculus |
organism.taxonId | integer | NCBI taxonomy identifier of the source organism; the hook for species-level slicing. | 9986 |
sequence.length | integer | Amino acid length of the canonical sequence. | 430 |
sequence.mass | integer | Computed molecular mass in daltons. | 2175 |
comments (cc_function) | text | Curated function annotation delivered with evidence codes, so each claim carries its reason to believe. | |
comments (cc_subcellular_location) | text | Subcellular location annotation with evidence attribution. | |
features (ft_domain, ft_signal, ft_disulfid, ft_mod_res) | string | Per-feature begin/end coordinates: domains, chains, signal peptides, transmembrane regions, disulfide bonds, modified residues. | |
go (go_id, go_p, go_c, go_f) | string | Gene Ontology term identifiers and names split by ontology branch - process, component, function. | GO:0005737 |
ec | string | Enzyme Commission numbers assigned to the protein, joining entries into the enzyme hierarchy. | 2.6.1.1 |
xref_pdb, xref_alphafolddb, xref_chembl, xref_drugbank, xref_reactome | string | Cross-reference columns reaching structure, chemistry, drug and pathway resources - a 300-plus pointer-field family. | |
entryAudit.firstPublicDate / lastAnnotationUpdateDate | date | Entry integration date and last annotation update date; the curation-history pair on every row. | 1989-10-01 / 2026-06-10 |
annotation_score | integer | Annotation completeness score from 0 to 5; usable as a label-confidence weight in models. | |
protein_existence | enum | Evidence level of protein existence, experimentally confirmed down to predicted-only. |
Questions buyers ask
How big is the UniProt rest api protein sequence function dataset?
149,810,139 UniProtKB entries in the current vintage - 575,503 manually reviewed Swiss-Prot records and 149,234,636 computationally annotated TrEMBL ones - backed by 381,103,551 UniRef clusters, 1,158,429,795 archived UniParc sequences and 1,086,758 organism proteomes.
What separates reviewed entries from computationally annotated ones?
Reviewed Swiss-Prot records carry curator-approved names, functions and evidence chains; the much larger computed layer extends the same eighteen-field schema to long-tail biology by automatic annotation. The entryType field keeps the tiers separable downstream, so any analysis can scope itself to the trust level it needs.
How far back do the protein records reach?
Entries have been integrated since 1986, and the oldest flat-file date in the verification capture reads 1 October 1989. Because every row indexes its own first-publication and last-annotation dates - entry P12345 shows 1989-10-01 against 2026-06-10 - historical states reconstruct cleanly, and vintage-stamped deliveries pin analyses to a citable snapshot.
What does one record contain?
One protein entry: accession and mnemonic identifier, reviewed-status tier, recommended name, gene symbols, organism and taxon ID, sequence length and computed mass, evidence-tagged function and subcellular location, feature coordinates, GO terms, EC numbers, cross-references, audit dates, an annotation-completeness score and a protein-existence tier - eighteen documented fields on a stable row shape.
Can UniProt data be joined against structure, chemistry and pathway corpora?
Yes. Stable accessions and NCBI taxon IDs carry the internal joins, more than 300 cross-reference columns reach resources such as PDB, AlphaFold DB, ChEMBL, DrugBank and Reactome, and the translation layer maps accessions against 100-plus external identifier types - so joins run on keys both sides already share.
Can a sample be scoped to my organisms or protein families?
Yes. Name the organisms, families, GO branches or annotation-score floor when you request the sample and it lands pre-cut on the same eighteen-field row shape - real rows in your pipeline before any commitment.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.