Biotechnology · KEGG (Kyoto Encyclopedia of Genes and Genomes)

KEGG

Datadory delivers biotechnology data covering KEGG, the Kyoto Encyclopedia of Genes and Genomes: 587 reference pathway maps, 28,430 KO ortholog groups, more than 67.8 million genes across 11,949 organisms, 19,623 metabolites, 12,473 reactions and 12,896 drug entries linking molecular parts to high-level functions. Delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
Global reference content spanning all domains of life plus viruses - organism, not country, is the axis that matters
How far back
Continuously maintained since 1995; current release notes dated July 1, 2026, with per-database update stamps through 2026-08-20
How fine
One structured entry per object - pathway map, ortholog group, gene, genome, compound, reaction, enzyme, drug or disease

What is KEGG data?

Sixteen databases wearing one name. KEGG - the Kyoto Encyclopedia of Genes and Genomes, maintained by Kanehisa Laboratories since 1995 - is the reference resource that turns molecular parts lists into statements about what biological systems actually do, and Datadory delivers its contents as typed rows rather than a browse interface.

The sixteen fall into four groups. Systems information: PATHWAY's 587 reference maps, BRITE functional hierarchies and MODULE functional units. Genomic information: the KO orthology table at 28,430 entries, with GENES holding more than 67.8 million genes across 11,949 organisms beneath per-organism GENOME records. Chemical information: COMPOUND's 19,623 metabolites, GLYCAN structures, 12,473 reactions and ENZYME nomenclature at 6,972 entries. Health information: NETWORK interaction maps, 3,098 DISEASE entries, 12,896 DRUG entries and MEDICUS.

That gene-to-function linkage is why KEGG remains the default citation for systems biology, metabolomics interpretation and biotech target mapping - the consensus layer that pure volume repositories lack. Within Datadory's catalog of 1,744 datasets across 159 viable industries, this record scores 8/10 against a 7.81 catalog mean, one of twenty-four biotechnology records. Get a sample of this dataset cut to your pathways and organisms.

What does a sample row of KEGG data look like?

One structured flat-file block per entry, plus the two-column list view and the statistics header that sizes each database:

# one structured entry per object - here, a human pathway map
ENTRY            hsa04010                    Pathway
NAME             MAPK signaling pathway - Homo sapiens (human)
CLASS            Environmental Information Processing; Signal transduction
PATHWAY_MAP      hsa04010  MAPK signaling pathway

# the identifier-list view - two columns, one row per entry
map01100         Metabolic pathways
map01110         Biosynthesis of secondary metabolites

# the statistics header that sizes each database
database=ko      entries=28,430     stamp=2026/08/20

# shape of the slice
grain=entry      types=pathway|ortholog|gene|genome|compound|reaction|enzyme|drug|disease

Read what the rows prove. The pathway block carries identity (ENTRY), a source-qualified name, a classification path (Environmental Information Processing; Signal transduction reads as category, then subcategory) and the map pointer that hangs drawing-level detail off the same identifier. The list view underneath is the shape most join keys take - one row, one identifier, one preferred name - and the statistics header is why scale claims stay checkable: the ko database states 28,430 entries against its own 2026/08/20 stamp rather than asking you to trust a brochure. Rows shown illustrate the verified record shape; a requested sample pulls live entries scoped to the pathways, organisms or molecules you name.

What fields does KEGG data include?

Fifteen documented fields ride on the entry structure, covering identity, classification, genomic placement, chemistry and literature provenance. Definitions below were verified against live records during the August 2026 research pass. Where a field's definition is documented but no worked example cell was captured in that pass, the example column stays empty rather than being decorated with an invented figure - and anything adjacent confirmed during sample preparation folds under additional fields on request rather than being promised blind.

What does coverage look like across geography, time and granularity?

Geography - the subject is life itself rather than any jurisdiction: reference content spans all domains of life plus viruses, and organism - not country - is the filter that matters. A kinase conserved from yeast to human carries one functional meaning across every organism code that shares it.

Temporal - continuously maintained since 1995, which makes KEGG one of the longest-lived curated resources in bioinformatics. The record dates itself: current release notes are dated July 1, 2026, and individual databases carry their own update stamps - the orthology table's most recent read 2026/08/20 - so any extract quotes a checkable currency figure instead of a vibe. Retaining successive deliveries lets you diff release over release rather than re-profiling anything yourself.

Granularity - one structured entry per object: pathway map, ortholog group, gene, genome, compound, reaction, enzyme, drug or disease. Scale check: 587 reference maps, 28,430 orthology groups, more than 67.8 million genes across 11,949 organisms, 19,623 metabolites, 12,473 reactions, 6,972 enzymes and 12,896 drug entries. The orthology layer is the quiet asset - attaching one functional meaning to genes from thousands of organisms took thirty years of curation, which is exactly what no amount of raw sequencing replicates.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel your stack already speaks. The rows arrive identical either way - entries parsed out of the flat files, identifiers resolved, cross-references kept attached - so nothing in your pipeline re-parses reference prose by hand.

Every delivery ships with the field dictionary above plus the sample rows for validation, so your first join attempt happens against evidence rather than hope.

Who uses this data, and for what?

  • Functional annotation of gene lists - a differential-expression hit list becomes a pathway story when every gene resolves onto its KO ortholog group and from there onto the maps it participates in; worked flows continue on our ML model training page.
  • Metabolomics interpretation - measured peaks resolve against 19,623 COMPOUND entries and their reaction links, so an annotated spectrum ends at named biochemistry instead of an unnamed mass.
  • Target and mechanism mapping - pathway membership plus enzyme classification puts candidate targets inside their mechanism context before a program commits; see market sizing for the downstream sizing workflow.
  • Knowledge-graph construction - gene-to-pathway-to-compound-to-drug edges arrive typed and keyed, ready to extend outward to sequence, structure and interaction resources on identifiers both sides already share.
  • Cross-species comparability - KO ortholog groups are organism-neutral by design, so a mechanism demonstrated in mouse maps onto human without fuzzy symbol matching.
  • Competitive and diligence work - which mechanisms have reference maps, and how densely drugs cluster on them, is readable straight from the health-information group.

Which personas get the most value?

Data scientists and ML engineers get a functional label for every gene - orthology membership, pathway placement, enzyme class - which turns identifier soup into supervised features; the fuller workflow lives on data scientists working in biotechnology. Bioinformaticians and computational biologists get thirty years of curated consensus behind their enrichment analyses, so a pathway claim survives review. Market researchers and consultants ground technology landscapes in accepted biology instead of vendor paraphrase. Developers building data products get a stable identifier space and a self-describing entry structure that loads once and joins outward cleanly; see developers building in biotechnology. Investors and diligence teams treat the citation footprint as platform moat. Journalists, academics and students get citable answers to 'what does this gene actually do'.

Which notes pair with this dataset?

Notes that pair well with this page:

  • Biotechnology data hub - the pooled industry view this record sits inside, alongside sequence, structure and study corpora.
  • KEGG DRUG Database - the health-information sibling: 12,896 approved-drug entries with targets, metabolism and classification codes.
  • The KEGG source profile - publisher scope, resource layout and dating behavior of the underlying estate.
  • UniProt protein sequence data - the sequence-and-function layer these records cross-reference on shared identifiers.
  • Human Protein Atlas - tissue-level protein evidence to test whether a mapped function shows up where it matters.
  • STRING protein interaction data - put each mapped gene back inside its interaction context.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - KEGG data; remainder folded below
fieldtypedefinitionexample
ENTRYstringEntry identifier and entry type opening every KEGG database record; the primary key all downstream systems anchor to.hsa04010 Pathway
NAMEstringEntry name, qualified by source vocabulary where applicable.MAPK signaling pathway - Homo sapiens (human)
DESCRIPTIONtextTextual summary of the entry (PATHWAY, DISEASE, MODULE, KO, REACTION, NETWORK).
CLASSstringClassification category of the entry - category and subcategory separated by a semicolon.Environmental Information Processing; Signal transduction
PATHWAY_MAPstringPathway map identifier and name for PATHWAY entries.hsa04010 MAPK signaling pathway
ORTHOLOGYstringAssociated KO (KEGG Orthology) groups giving the entry its organism-neutral functional identity.
GENEStextOrganism-specific genes linked to a KO, pathway or genome entry.
ORGANISMstringOrganism code and name (GENES, GENOME entries).
POSITIONstringGenomic location of a gene (GENES entries).
AASEQ / NTSEQtextAmino-acid or nucleotide sequence carried on a gene entry (GENES).
FORMULAstringChemical formula of a COMPOUND entry.C19H15O4. Na
REACTION / EQUATIONtextBiochemical reaction equation (REACTION entries) or associated reactions (ENZYME).
DBLINKStextCross-references to external databases such as NCBI, UniProt and PubChem.CAS: 129-06-6
REFERENCEtextLiterature citations with AUTHORS, TITLE and JOURNAL sub-fields.
TARGETtextDrug target information on DRUG entries, keyed to human gene and KO identifiers.VKORC1 [HSA:79001] [KO:K05357]
Additional fields on request-GENOME-record extensions (org code, taxonomy lineage, statistics), exact mass and molecular weight companions on compounds, full sequence retrievals, MOL/KCF structure renditions, KGML pathway markup, JSON BRITE hierarchies, RDF graph serializations, drug-interaction pair tables and DGROUP groupings.on request

What teams do with it

  • Functional annotation of gene lists A differential-expression hit list becomes a pathway story when every gene resolves onto its KO ortholog group and from there onto the maps it participates in.
  • Metabolomics interpretation Measured peaks resolve against 19,623 COMPOUND entries and their reaction links, so an annotated spectrum ends at named biochemistry instead of an unnamed mass.
  • Target and mechanism mapping Pathway membership plus enzyme classification puts candidate targets inside their mechanism context before any program commits to one.
  • Knowledge-graph construction Gene-to-pathway-to-compound-to-drug edges arrive typed and keyed, ready to extend outward to sequence, structure and interaction resources both sides already share.
  • Cross-species comparability KO ortholog groups are organism-neutral by design, so a claim demonstrated in mouse maps onto human without fuzzy gene-symbol matching.

Questions buyers ask

How many organisms and genes does KEGG cover?

GENES holds more than 67.8 million genes across 11,949 organisms, spanning all domains of life plus viruses. Above it sits the KO orthology table at 28,430 entries, which attaches functional meaning - pathway membership, enzyme class - to organism-specific genes so a function learned in one species travels to the rest.

What are the four groups of KEGG databases?

Systems information - 587 PATHWAY reference maps, BRITE hierarchies and MODULE units; genomic information - KO orthologs, GENES and GENOME; chemical information - COMPOUND metabolites, GLYCAN, reactions and ENZYME nomenclature; and health information - NETWORK, DISEASE, DRUG and MEDICUS. One identifier space runs across all sixteen.

How current is KEGG data?

Continuously maintained since 1995, and the record dates itself: release notes are dated July 1, 2026, with per-database stamps such as the orthology table's 2026/08/20. Deliveries run daily, weekly, or hourly on your side, and retained successive deliveries make release-over-release diffing routine rather than heroic.

What can KEGG data be joined against?

Every entry carries DBLINKS cross-references, and the identifier space converts directly into NCBI GeneID, NCBI ProteinID, UniProt, PubChem and ChEBI namespaces. Genes therefore reach sequence archives, compounds reach chemistry collections and drugs reach approval registries without a separate mapping project sitting in front of the analysis.

Does KEGG cover drugs and diseases too?

Yes - the health-information group holds 12,896 DRUG entries and 3,098 DISEASE entries beside NETWORK and MEDICUS. The DRUG estate is deep as well as wide: 7,335 entries carry therapeutic target annotations, 1,395 list metabolizing enzymes and 2,550 tie to Japanese label codes with 2,167 into FDA labels.

Why do pathway maps matter when repositories hold more raw data?

Because maps are the interpretation layer, not the volume layer. UniProt, PubChem and the sequence archives dwarf KEGG on raw counts; what those resources lack is a curated consensus on what the parts do together. Enrichment analysis, metabolomics interpretation and mechanism mapping all need that consensus, and 587 reference maps are where it lives.

Can a sample be scoped to specific pathways or organisms?

Yes. Name the maps, ortholog groups, organisms or molecule classes you care about - or describe them and the scope gets resolved against the current release - and the extract returns shaped exactly to that slice with the field dictionary and cross-references attached before any recurring delivery begins.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing