Human Protein Atlas

Datadory delivers biotechnology data covering Human Protein Atlas: a Swedish research program that maps the spatial expression of essentially every human protein-coding gene - 27,883 antibodies raised against 17,407 proteins, tissue-level immunohistochemistry across 45 normal tissue types and 20 cancer types, transcriptomics across 51 tissue types, single-cell profiles spanning 34 tissues and 154 cell types, and subcellular localization resolved into 49 organelle classes. Delivered as API, files, or your warehouse, on the cadence you choose.

What is the Human Protein Atlas?

The Human Protein Atlas is a Swedish research program, based at SciLifeLab and designated a Global Core Biodata Resource, with one obsession: where in the human body does each protein actually appear. Its answer spans nine linked resources - Tissue, Brain, Single Cell, Subcellular, Cancer, Blood, Cell Line, Structure and Interaction - built on 27,883 antibodies raised against 17,407 unique proteins.

Three complementary measurement styles carry the map. Antibody-based immunohistochemistry images actual stained tissue sections across 45 normal tissue types and 20 cancer types, reaching 15,312 genes with tissue staining. Transcriptomics adds quantitative RNA depth across 51 tissue types. Mass-spectrometry proteomics confirms the protein side. Around those sit single-cell resolution over 34 tissue types and 154 cell types, subcellular localization resolving 13,603 genes into 49 organelle classes, blood plasma profiling, and predicted structures for 19,904 proteins.

The consequence for anyone choosing drug targets, designing assays or writing about human biology: tissue restriction, subcellular address and cancer behavior become lookups rather than literature reviews. Get a sample of this dataset and inspect real records before committing pipeline time.

What does a sample row look like?

One row per gene, with identity, classification, expression verdicts and validation tier side by side. Cyclin B1, a cell-cycle regulator any oncology team will recognize:

Gene                     CCNB1
Gene synonym             CCNB
Ensembl                  ENSG00000134057
Gene description         Cyclin B1
Uniprot                  P14635
Chromosome               5
Position                 69167135-69178245
Protein class            Cancer-related genes | Essential proteins | Plasma proteins
RNA tissue specificity   Tissue enhanced
RNA tissue nTPM          bone marrow: 49.9 | lymphoid tissue: 70.9
Single cell specificity  Cell type enhanced
Reliability (IH)         Enhanced
Subcellular location     Cytosol

Read it left to right and you have everything a triage meeting argues about: stable identifiers on both the gene and protein side, genomic coordinates, functional classification, the transcriptomic verdict (enhanced in bone marrow and lymphoid tissue, and which cell types drive it), how much to trust the antibody staining itself, and where in the cell the protein sits. The flat row is the index; behind it hang per-tissue, per-cancer-type, per-cell-line and per-cell-type expression matrices, plus the staining imagery each verdict summarizes.

What fields does the Human Protein Atlas include?

Seventeen core fields travel on every protein record, covering identity, function, transcriptomic specificity and antibody-validation status. Definitions below come from the documented schema; entries flagged in the footnote are folded into your sample on request.

What does coverage look like across geography, time and granularity?

Geography - the subject is human biology rather than any jurisdiction: samples profile human tissues, cells and cancers, with comparative pig and mouse brain sections included alongside the human brain resource for cross-species context.

Temporal - organized as versioned releases dating back to 2005, each a complete citable snapshot; version 25.1 arrived in May 2026, built on Ensembl 109 gene annotation. Pin analyses to a release number and they stay reproducible, and retaining successive deliveries lets you diff release over release instead of re-profiling anything yourself.

Granularity - one record per gene/protein, with expression values hanging off it per tissue, per cell type, per cancer type and per cell line. A question like 'which kinases are detectable in heart tissue but absent from liver' is a filter over those matrices, not a re-read of papers.

Scale check: 27,883 antibodies, 17,407 proteins, 15,312 genes with tissue immunohistochemistry, 13,603 genes with subcellular localization, 154 cell types, 49 organelle classes and 19,904 predicted structures. The antibody fleet is the scarce asset - generating validated, tissue-stained evidence at this width took a decade-scale program, which is why the corpus has no practical substitute.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Who uses this data, and for what?

  • Target prioritization and tissue-restriction screens - rank candidates whose protein appears in the tissue of interest but stays quiet everywhere else, using RNA tissue specificity categories backed by staining evidence rather than a single transcriptomics study.
  • Safety and off-target assessment - flip the same query: RNA tissue distribution flags anything detected broadly, so pan-tissue liabilities surface before they reach an animal study.
  • Antibody and assay development - the Reliability (IH) tier on each stain tells assay teams which targets arrive with enhanced validation and which need their own confirmatory work planned in.
  • Cancer biomarker discovery - cancer-resource specificity separates markers enriched in tumor tissue from general housekeeping signal, with 20 cancer types covered side by side.
  • Single-cell deconvolution - cell-type-level specificity across 154 cell types lets analysts decompose bulk tissue profiles into the cell populations that produced them.
  • Knowledge-graph construction - Ensembl identifiers and UniProt accessions give every node keys that resolve deterministically into sequence, structure and interaction resources both sides already share.

Which personas get the most value?

Bioinformaticians and computational biologists get staining evidence and transcriptomics reconciled onto one identifier per gene, so the protein-level claim and the RNA-level claim stop living in different documents. Data scientists and ML engineers get labeled specificity categories and numeric normalized-expression values - supervision for tissue-of-origin classifiers and target-prioritization models. Competitive-intel and product teams in biotech read which targets already carry enhanced-validated staining when judging how crowded a mechanism actually is. Developers building data products get a flat, keyed record shape that joins outward on Ensembl and UniProt without fuzzy matching. Investors and diligence teams check whether a claimed tissue-restricted target really is restricted before the term sheet does. Journalists, academics and students get citable, image-backed answers to 'where does this protein appear in the body'.

Which notes pair with this dataset?

Notes that pair well with this page:

  • Biotechnology data hub - the pooled industry view this record sits inside, alongside sequence, structure and interaction corpora.
  • Human Protein Atlas vs NCBI GEO - curated spatial staining evidence against raw submitted expression series; the pair brackets curated versus raw.
  • The Human Protein Atlas source profile - program scope, resource layout and versioning behavior of the underlying atlas.
  • UniProt protein sequence data - the functional annotation layer these records cross-reference on UniProt accessions.
  • STRING protein interaction data - put each localized protein back into its interaction context.
  • RCSB Protein Data Bank structures - experimental structures beside the atlas-published predicted models.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - seventeen core fields on every protein record; remainder folded below
fieldtypedefinitionexample
GenestringHGNC gene symbol.CCNB1
Gene synonymstringAlternative gene symbols.CCNB
EnsemblstringEnsembl gene identifier - the primary key joining records to external gene resources.ENSG00000134057
Gene descriptionstringGene name description.Cyclin B1
UniprotstringUniProt protein accession cross-reference.P14635
ChromosomestringChromosome assignment of the gene.5
PositionstringGenomic coordinate range of the gene.69167135-69178245
Protein classtextComma-separated classification terms such as Cancer-related genes, Plasma proteins, Transporters.Cancer-related genes, Plasma proteins
Biological processtextAssociated Gene Ontology biological process terms.
Molecular functiontextAssociated Gene Ontology molecular function terms.
Disease involvementtextDisease relevance annotation.
EvidenceenumOverall protein evidence level for the gene.Evidence at protein level
RNA tissue specificityenumSpecificity category of tissue expression: Tissue enriched, Group enhanced, Low specificity and kin.Tissue enhanced
RNA tissue distributionenumBreadth of tissue detection, e.g. Detected in all, Detected in many.Detected in many
RNA tissue specific nTPMtextNormalized TPM values for the most specific tissues, as tissue:value pairs.bone marrow: 49.9; lymphoid tissue: 70.9
RNA single cell type specificityenumSpecificity category across single-cell clusters.Cell type enhanced
Reliability (IH)enumImmunohistochemistry antibody validation tier: Enhanced, Supported, Approved or Uncertain.Enhanced

Questions buyers ask

How big is the Human Protein Atlas?

The current generation rests on 27,883 antibodies targeting 17,407 unique proteins. Tissue-level immunohistochemistry covers 15,312 genes across 45 normal tissue types and 20 cancer types, subcellular localization resolves 13,603 genes into 49 organelle classes, single-cell data spans 34 tissue types and 154 cell types, and predicted structures exist for 19,904 proteins.

Does the atlas cover anything besides healthy tissue?

Yes. Alongside the normal-tissue resource sit dedicated Cancer (20 cancer types), Cell Line, Blood (plasma profiling), Brain (including comparative pig and mouse sections), Structure and Interaction resources. A target can be checked against tumor behavior, cultured-cell behavior and circulating-blood presence without leaving the same identifier space.

How does antibody-based evidence differ from the RNA evidence?

They measure different layers and corroborate each other. Immunohistochemistry shows the actual protein in stained tissue sections, graded by a validation tier - Enhanced, Supported, Approved or Uncertain - that states how strongly the antibody was validated. Transcriptomics contributes quantitative normalized expression values and specificity categories such as Tissue enriched or Low specificity. Disagreement between the two layers is itself informative: RNA present without detectable protein usually means post-transcriptional control.

How far back does the data go?

Versioned releases date to 2005, each a complete citable snapshot. Version 25.1 was published in May 2026, built on Ensembl 109 gene annotation. Analyses pinned to a release number stay reproducible, and successive releases can be diffed to see how coverage widened.

What can Human Protein Atlas data be joined against?

Every record carries an Ensembl gene identifier, a UniProt accession and HGNC symbols, which connect it deterministically to sequence databases, structure repositories, interaction networks and expression-series archives. Genomic coordinates and GO terms extend the same joins into annotation pipelines and ontology-aware tools.

Is single-cell resolution available for every tissue and gene?

Not uniformly, and the fields make the boundary explicit. Single-cell data covers 34 tissue types and 154 cell types, while whole-tissue transcriptomics covers 51 tissue types - so a gene record may carry tissue-level specificity with cell-type detail still pending. Filtering on the specificity fields themselves keeps claims scoped to the resolution that actually supports them.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing