Datadory notebook

Protein-protein interaction network databases: where the edges live

Datadory delivers biotechnology data covering protein-protein interaction networks at full depth: STRING's 20-billion-plus scored associations across 59.3 million proteins in 12,535 organisms, every pair carrying a combined confidence score from 0 to 1 decomposed into seven evidence-channel subscores, joined downstream by GWAS loci, tissue expression, proteomic identifications and measured bioactivity. Typed rows delivered daily, weekly, or hourly - your call.

1,744 datasets. Pick your catch.

What is a protein-protein interaction network database?

A protein-protein interaction network treats proteins as nodes and scored relationships as edges, so a gene list becomes a graph you can walk, cluster and rank. The reusable artifact is not a pathway diagram - it is a per-pair table where every row carries a confidence number and, in the best databases, a breakdown by evidence channel.

STRING Protein Interaction Database is the canonical example in Datadory's biotechnology catalog: more than 20 billion interactions across 59.3 million proteins, spanning 12,535 organisms from bacteria to animals. Granularity is per protein-pair interaction with per-channel subscores, plus per-protein accessory records. That structure is what separates a scored network from a list of co-citations - you can threshold, weight and trace exactly why two proteins sit next to each other.

Scale is the headline number. Of the 24 primary biotechnology datasets Datadory catalogs, none approaches STRING's edge count; the nearest neighbors in ambition are UniProt's 149,810,139 protein entries and PubChem's ~179.5 million compound records, both of which describe single entities rather than relationships between them.

How stable is the network once you pin a version?

A scored network is only useful if it reproduces, and STRING's release model is built for exactly that. The corpus ships as versioned releases running back past v3.0 before 2003 up to version 12.0, live since 26 July 2023, and each release carries its own scoring pass - so a pinned version yields identical edge weights across runs, notebooks and collaborators.

That stability is what makes the network a scaffold rather than a snapshot. Pin v12.0 and a target-prioritization result survives review, because whoever checks the work re-runs against the same graph instead of a quietly newer one. The layers around the network accumulate independently - expression cohorts across NCBI GEO's 143,551 RNA-seq Series, proteomic identifications across PRIDE's 40,809 projects, tissue and subcellular context across Human Protein Atlas's 17,407 profiled proteins - so treat those as additive columns joined onto a fixed edge set rather than as reasons to rebuild the graph.

The practical pattern: freeze the edge set at a named release, join the accumulating layers by stable protein identifiers, and diff network growth release-over-release when a new version lands. Retaining successive deliveries turns 'did the biology change or the database?' into a two-column comparison instead of an archaeology dig.

How is protein interaction network data delivered?

Name the organism, score floor and channels you care about when you request a sample and the extract arrives cut to that scope; the standing feed follows the same shape, so anything prototyped on the sample survives delivery intact. The record's field-by-field breakdown lives on the STRING Protein Interaction Database dataset page, alongside captured sample rows.

What does an interaction score actually mean?

Every STRING edge carries a confidence score built from per-channel subscores - text mining, co-expression, and the other evidence channels the database aggregates. Reading the subscores, not just the composite, is what keeps a network honest: two proteins can reach the same composite score through a single high-throughput assay or through converging independent evidence, and only the subscore breakdown tells you which.

Three usage rules follow. First, threshold by channel when your question is mechanistic - a text-mining-heavy edge is a hypothesis about the literature, not about the cell. Second, keep the score with the edge wherever the data travels; every delivered row carries the composite and its seven subscores, so a filtered network stays auditable. Third, re-threshold when you swap organisms, because evidence density is not uniform across the 12,535 species covered.

Which datasets turn a network into a target list?

A network alone ranks connectivity; a target list needs genetics, expression and chemistry attached. The biotechnology slice of Datadory's catalog supplies each layer.

LayerDatasetWhat it contributesScale
Scored edgesSTRING Protein Interaction DatabasePer-pair confidence with per-channel subscores20+ billion interactions, 59.3 million proteins, 12,535 organisms
Genetic priorityGWAS CatalogCurated SNP-trait associations with EFO mapping229,900 studies; 1,191,572 associations
Expression filterHuman Protein AtlasTissue, cell-type and cancer expression27,883 antibodies, 17,407 proteins
Proteomic evidencePRIDE Proteomics ArchiveRaw and processed mass-spectrometry identifications40,809 projects; 500-900 new submissions per month
Pathway mappingKEGGCurated pathway maps and ortholog groups587 pathway maps; 28,430 KO entries; 67.8M genes
DruggabilityChEMBLCurated bioactivity against targets24,527,044 activities on 18,552 targets

The join logic runs downstream from the network: GWAS Catalog associations nominate loci, STRING resolves which proteins inside a locus actually talk to each other, Human Protein Atlas checks tissue expression, and ChEMBL reports whether anyone has already measured activity against the surviving candidates. ChEMBL's 2.92 million documented molecules make that last step a lookup rather than a literature search.

Who builds on a scored interaction network?

A scored network is infrastructure, and the people who lean on it divide cleanly by what they do with the edges.

Bioinformaticians and computational biologists get every pair decomposed channel by channel, so a network claim survives handoff into someone else's pipeline with its evidence attached. Data scientists and ML engineers train on labeled edges - per-channel supervision plus per-protein context turns link prediction and target ranking into a supervised problem instead of a heuristic one. Target-prioritization teams in biotech intersect the graph with genetic loci and expression filters to collapse a wide association list into a shortlist whose members demonstrably talk to each other. Competitive-intel and product teams read neighborhood density around a mechanism before judging how crowded a target space really is. Investors and diligence teams check whether a platform's claimed network effects actually surface in the scored edges before the term sheet goes out. Journalists, academics and students get citable answers to the question behind most coverage - which proteins does this one travel with?

Across all of them the pattern repeats: the network supplies the shape, the surrounding datasets supply the constraints, and the deliverable is a decision with evidence attached rather than another diagram.

Which caveats change the conclusions you draw?

Three caveats decide whether a network conclusion holds.

First, a composite score is a summary, not a verdict. Two pairs can share a 0.700 composite while one earned it at the bench and the other from co-mentioned abstracts; only the channel subscores distinguish them. Threshold on the channels your question actually trusts before quoting any headline number.

Second, evidence density tracks how thoroughly an organism has been studied. Model organisms carry dense experimental and curated-pathway support, while the long tail of the 12,535 covered species leans harder on transferred evidence - conserved gene neighborhood, fusion events, phylogenetic cooccurrence, and observations transferred along orthology from better-studied relatives. Cross-species conclusions inherit that asymmetry, so re-threshold whenever you swap organisms.

Third, identifiers are the join risk, not the scores. Protein designations drift across databases and eras; normalize through a stable protein namespace before resolving edges, or a filtered network silently loses pairs to unmatched keys rather than to biology. None of this argues against the network - it argues for keeping the subscores, the taxonomy ID and a normalized identifier on every row you ship, which is precisely how the records arrive.

Where does a PPI network fit in a biotech data stack?

Interaction networks sit in the middle of the stack: upstream of them are sequences and genomics, downstream are pathways, structures and assays. Datadory catalogs 27 datasets for biotechnology (24 primary plus 3 related) spanning sequence archives, structure resources, chemistry, registries and regulatory records, and the record shapes across the slice range from flat protein entries to raw reads and molecular structures - every one of which feeds a network workflow somewhere.

Upstream, Ensembl Genome Browser (release 116, June 2026, roughly 60k human genes) supplies the coordinates that map a locus to its proteins, and NCBI SRA holds the raw reads at tens of petabases for reanalysis. Downstream, KEGG's 587 pathway maps give the networks a functional vocabulary, and ChEMBL's assay records test whether network-identified candidates have measurable biology.

The connective tissue is identifiers. STRING keys on stable protein identifiers and UniProt normalizes across naming schemes, which is why the pair anchors most network pipelines. Start from a gene list, normalize through UniProt, resolve edges in STRING, and every downstream dataset in the catalog joins cleanly from there.

Where to go next

The biotechnology data guide covers all 27 cataloged datasets for the industry. For the genetic half of target prioritization, the genome-wide association study guide walks from SNP-trait associations to a ranked candidate list. For the structural half, the compound bioactivity screening guide shows how ChEMBL and PubChem records validate a network-derived candidate against measured assay data.

Go straight to the records: the STRING Protein Interaction Database, the GWAS Catalog, Human Protein Atlas and ChEMBL. The wider view lives on the biotechnology data hub and in the best biotechnology datasets ranking. Request a sample scoped to your organisms, score floor and channels - real scored pairs come back before anything recurring switches on.

Datasets that complete a protein interaction network workflow (Datadory catalog, as of August 2026)
layerdatasetwhat it contributesscale
Scored edgesSTRING Protein Interaction DatabasePer-pair confidence with per-channel subscores20+ billion interactions, 59.3 million proteins, 12,535 organisms
Genetic priorityGWAS CatalogCurated SNP-trait associations with EFO trait mapping229,900 studies; 1,191,572 associations
Expression filterHuman Protein AtlasTissue, cell-type and cancer expression plus subcellular localization27,883 antibodies against 17,407 proteins
Proteomic evidencePRIDE Proteomics ArchiveRaw and processed mass-spectrometry identifications40,809 projects; 500-900 new submissions per month
Pathway mappingKEGGCurated pathway maps and ortholog groups587 pathway maps; 28,430 KO entries; 67.8M genes
DruggabilityChEMBLCurated bioactivity records against protein targets24,527,044 activities on 18,552 targets

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Biotechnology Global

STRING Protein Interaction Database

Biotechnology

UniProt

Biotechnology Global submissions

NCBI GEO

Biotechnology Global - laboratory submissions worldwide via ProteomeXchange…

PRIDE Proteomics Archive data

accession · doi · submissionDate …+12 more

Biotechnology

Human Protein Atlas

Biotechnology

RCSB Protein Data Bank

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

What is the best protein-protein interaction network database?

STRING Protein Interaction Database is the reference source. It scores more than 20 billion interactions across 59.3 million proteins in 12,535 organisms, with version 12.0 live since 26 July 2023 and the release line running back past v3.0 before 2003. Its per-channel subscores let you keep only the evidence types your analysis trusts.

How big is the STRING interaction network?

More than 20 billion scored protein-protein associations spanning 59.3 million proteins across bacteria, archaea, plants, fungi and animals. One record per protein pair carries the combined 0-to-1 confidence score plus seven channel subscores, with accessory per-protein records holding identity and functional annotation around the edges. The corpus runs to hundreds of gigabytes before channel-level detail is even unpacked, which is why nobody rebuilds it locally from scratch.

Can you separate experimentally backed edges from predicted ones?

Yes - that is what the channel decomposition is for. The experimental subscore consolidates biochemical and biophysical evidence from sources such as BioGRID, IntAct and PDB, while the coexpression, text-mining, neighborhood, fusion and cooccurrence subscores carry the predicted channels. Filter on the experimental channel alone for bench-grounded subnetworks, or weight channels explicitly when a mechanistic claim needs converging independent evidence rather than a single assay.

How do I turn an interaction network into a pathway result?

Score your gene list against STRING's network to surface connected clusters, then map those clusters onto curated pathway maps. KEGG publishes 587 pathway maps with 28,430 KO entries across 11,949 organisms. Pair the network edges with KEGG's maps and you move from a protein list to named pathways with evidence attached.

Which datasets add expression and structure context to an interaction network?

Human Protein Atlas contributes tissue and subcellular context for 17,407 unique proteins using 27,883 antibodies, and RCSB Protein Data Bank adds 3D structure for about 258,616 experimental entries plus more than 1 million computed models, so filtering an interaction network by where and in what shape its proteins exist is a join rather than a project.