Life & Health Insurance · BRENDA
BRENDA Enzyme Database
Datadory delivers brenda enzyme database data covering roughly 8,831 EC classes with about 40,894 KM values, some 94,073 turnover numbers, catalytic-efficiency ratios, Ki and IC50 inhibition constants, cofactors, organisms and over 32 million linked sequence records - each measurement keyed to enzyme, organism and cited reference. Name your classes and get real rows first.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Global by biology rather than by market: organisms across all kingdoms of life, from thermophilic archaea to human tissue sources. The geography that matters here is taxonomic - the same reaction measured across species sits side by side, so 'whose version of this enzyme is fastest' resolves as a filter on the Organism column instead of a review of whoever published last.
- How far back
- Continuously curated since 1987, with numbered releases marking milestones - Release 2026.1, dated March 4, 2026, added 124 new and updated 750 enzyme classes. The resource's own statistics series reaches back to 07.2007, so coverage growth itself is chartable. Longitudinal series accrue from repeated capture on your cadence, which is why teams start one before they need the trend.
- How fine
- One entry per enzyme-organism-(protein-)commentary-reference level per field. That is the whole trick of this dataset: a KM value never floats free of its organism, its assay commentary and its cited paper. Averages computed downstream stay reproducible because the atomic measurement - not a summary table - is what ships.
What is the BRENDA Enzyme Database dataset?
It is the measured half of the lab bench, delivered as rows instead of browser tabs. Where the rest of our life health insurance catalog tabulates deaths, premiums and claims, the BRENDA Enzyme Database - curated since 1987 at the Leibniz Institute DSMZ, an ELIXIR Core Data Resource and a Global Core Biodata Resource - tabulates biochemistry itself: roughly 8,831 EC classes carrying about 40,894 KM values, some 94,073 turnover numbers, catalytic-efficiency ratios and inhibition constants.
The grain does the heavy lifting. Entries exist at enzyme-organism-commentary-reference level, so a single EC class holds separate measured values per organism, tissue and literature source - invertase under EC 3.2.1.26 arrives with its Michaelis constant, its reaction equation and commentary explaining the assay conditions, not just a name. Around the kinetics sit cofactors and metal ions (about 40,907 activating-compound entries), molecular weights, subunit compositions, pH and temperature optima, inhibitor and activator lists, and cross-references into more than 32 million linked sequence records.
Within the Datadory catalog this slice scores 9/10 - top-tier among the 21 datasets pooled under Life & Health Insurance, against a 7.81 average across all 1,744 datasets we catalog. Get a sample of this dataset and we return rows shaped exactly like the dictionary below, cut to the enzyme classes and organisms you name.
What do sample rows from this dataset look like?
Each delivery lands as one kinetic entry per enzyme-organism-commentary-reference combination, banded exactly like this:
# delivered grain: one kinetic entry per enzyme x organism x commentary x reference
EC Number : 3.2.1.26
Recommended Name : Invertase
Systematic Name : beta-D-fructofuranoside fructohydrolase
Synonyms : <alternate names, incl. beta-fructofuranosidase>
Reaction : sucrose + H2O = D-glucose + D-fructose
CAS RN : <registry number>
Organism / Source Tissue : Homo sapiens / <tissue>
Localization : <cellular compartment>
# kinetics band - measured numbers, each with its own commentary + literature reference
KM Value [mM] : e.g. 0.05
Turnover Number [1/s] : e.g. 120
kcat/KM [mM^-1 s^-1] : <catalytic-efficiency ratio>
Specific Activity [U/mg] : <where reported>
Ki Value / IC50 Value [mM] : <inhibition constant against a named inhibitor>
# molecular band
Sequence (UniProt Acc.) : <cross-reference> - 32M+ sequence links pool-wide
PDB ID : <structure cross-reference>
Cofactor / Metals Ions : <required coenzymes and metal ions>
Molecular Weight / Subunits: <observed mass and quaternary structure>
# condition band - operating windows
pH Optimum / Temperature Optimum / Stability
Solvent Stability / Storage Stability
# application band
Application / Cloned / Crystallization / Engineering-Mutations / DiseaseRead the anatomy rather than the placeholders. One entry hands over four things at once: the catalyst's identity and reaction chemistry, its measured performance numbers each carrying their own commentary and citation, the molecule and organism behind the number via sequence and structure cross-references, and the operating conditions plus applied annotations that decide whether the catalyst survives contact with a process. The two example values shown - a KM of 0.05 mM and a turnover number of 120 1/s - are the kind of concrete figures that turn how fast is this enzyme, really from a footnote chase into a cell in a spreadsheet.
Your sample doubles as validation: the anatomy above is the shipped schema, confirmed against live records for the enzyme classes you name, with per-class density checked before anything ships.
Which fields does the field dictionary define?
Thirteen field families, split across the six bands described above. The join key is the EC number: it travels across organisms, papers and internal systems, which is why cross-source reconciliation starts there instead of with enzyme names that drift between synonyms.
Three families do the heavy lifting. KM Value and Turnover Number arrive as measured points with commentary and references attached - roughly 40,894 and 94,073 entries respectively - so affinity and speed stay separable instead of collapsed into one adjective. kcat/KM pre-computes the efficiency ranking biocatalyst comparisons actually turn on. Ki Value / IC50 Value carries inhibition constants against named compounds, which is what makes a sensitivity screen a filter operation on delivered rows rather than a parsing project over prose.
Deeper structures - pathway context, activator lists, subunit-composition detail - fold under additional fields on request, pinned with examples once a sample names the families it needs.
Where does coverage run, and at what grain?
Three chips summarize the footprint:
- Geography: global by taxonomy, not by market - organisms across all kingdoms of life. Cross-species comparison inside one EC class is a group-by on the Organism column.
- Time frame: continuously curated since 1987, with numbered releases marking milestones - Release 2026.1 (March 4, 2026) added 124 new and updated 750 enzyme classes. The resource's statistics series reaching back to 07.2007 means coverage growth itself is chartable; longitudinal series accrue from repeated capture on whatever cadence you set.
- Granularity: one entry per enzyme-organism-commentary-reference per field. No averaging away the organism-by-organism spread - summaries average it away, the rows keep it.
Against the wider Datadory catalog - 1,744 datasets, average quality 7.81 - this slice holds a 9/10, earned by verified field documentation and the cleanest join spine in the category.
How is the data delivered through Datadory?
API, files, or your warehouse. Daily, weekly, or hourly.
Pick the channel your team already works in and set the cadence to match the decision being fed. Every delivery arrives keyed on EC number and organism, so each pull joins cleanly to the last and to whatever target, compound or organism lists you already hold. Field definitions travel with the data, multi-value kinetics stay parseable, and schema movement between curation passes gets flagged rather than discovered mid-pipeline.
Every delivery ships the field dictionary unchanged plus sample rows for validation, with records flattened to one observation per enzyme-organism-reference - so a cross-species comparison is a group-by rather than a merge-key archaeology project. Get a sample and the anatomy above comes back populated with your enzyme classes.
Who builds on this data, and for what?
Ranked by how directly a single kinetic entry settles the day job:
- Biocatalyst ranking and selection. kcat/KM efficiency ratios across candidate hosts turn 'which organism's version of this enzyme is fastest' into a sorted column instead of a literature review.
- Feature engineering and ML. Tens of thousands of labeled kinetic values with organism and mutation context assemble into parameter predictors without a labeling pass - patterns on the ml model training page.
- Competitive intelligence. Shifts in which EC classes attract characterization read as portfolio signals while rivals still quote last year's landscape - see the competitor tracking pattern.
- Inhibitor sensitivity screens. Ki and IC50 constants against named compounds land as columns, so a compound panel screens itself on arrival.
- Process design. pH optimum, temperature optimum, solvent tolerance and storage stability hold the operating windows that decide whether a catalyst survives scale-up.
- Target-validation signals for investors. Coverage growth per enzyme class is chartable back to 07.2007 - a quiet tell ahead of news flow; see citation-grade research for the sourcing discipline behind it.
For contrast inside the same pool: the BacDive strain metadatabase describes organisms thoroughly - taxonomy, culture conditions, biosafety - but measures no catalysis. The trade-offs are set out in the head-to-head comparison.
Which personas get the most value?
Data Scientists & ML Engineers (relevance 3 of 3) get labeled kinetic measurements with mutation context ready for model work - see data scientists in life & health insurance. Competitive Intelligence & Product Teams (3 of 3) read characterization effort as a portfolio signal. Market Researchers & Consultants (3 of 3) map biochemical coverage across organisms for landscape work - see market researchers use cases. Journalists, Academics & Students (2 of 3) cite the standard curated reference rather than someone's chart - see journalists academics use cases. Sales & Growth Teams (1 of 3) aim campaigns at labs studying named enzymes; Investors & Quant Researchers (1 of 3) read coverage growth as validation tells; Developers & Data-Product Builders (1 of 3) build lookup products on the EC-number spine.
What should I know before requesting a sample?
Three things, stated up front.
First, depth follows published research. A heavily studied enzyme class carries hundreds of kinetic entries; a recently classified one may carry a handful. Per-class density gets checked before anything ships rather than assumed uniform, so scope expectations by class, not by the headline totals.
Second, this is measured biochemistry, not strain metadata. No growth recipes, no biosafety levels, no culture conditions - those live in the BacDive strain metadatabase. If the question names an organism, start there; if it names a reaction, start here. Many teams need both halves in one deliverable.
Third, the join key discipline pays off downstream. Entries reconcile on EC number plus organism; samples cut to a named class-and-organism list come back with the field dictionary pinned against the delivered rows, which is also where additional field families lock for your pipeline.
Which pages pair with this dataset?
Notes worth reading next:
- Best life sciences tools & services datasets - where this record sits among the eight ranked, and what it beat.
- BacDive strain metadatabase - the descriptive counterpart: 102,187 strains with husbandry and safety, zero catalysis numbers (how it compares).
- ATCC Biological Materials Catalog - the acquisition-side pair: authenticated materials once the catalyst screen picks a source organism.
- BRENDA source profile - who curates the resource and how the record sits in our cataloging rubric.
- Life & health insurance data hub - the full pooled industry view, from mortality tables to MEPS microdata and this biochemical outlier.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
EC Number | string | IUBMB enzyme class identifier organizing the whole resource - the join spine every other table lands on. | 3.2.1.26 |
Recommended Name / Systematic Name / Synonyms | string | Nomenclature triple per class: the working name, the IUBMB systematic name and alternate names, so renamed or colloquial enzymes still resolve. | Invertase |
Reaction / Substrates Products / CAS RN | string | The reaction equation with substrate and product lists and the Chemical Abstracts registry number - the chemistry identity behind the class. | sucrose + H2O = D-glucose + D-fructose |
KM Value [mM] | number | Michaelis constant per enzyme-organism-commentary-reference entry - substrate affinity as a measured number, not a range in prose. Roughly 40,894 entries. | 0.05 |
Turnover Number [1/s] | number | kcat per entry, with commentary explaining assay conditions and the cited reference attached. Some 94,073 entries. | 120 |
kcat/KM [mM^-1 s^-1] | number | Catalytic-efficiency ratio - the ranking number when candidate biocatalysts compete. | - |
Ki Value / IC50 Value [mM] | number | Inhibition constants and half-maximal inhibitory concentrations tied to named inhibitors - the sensitivity screen in column form. | - |
Specific Activity [U/mg] | number | Activity per mass of enzyme preparation where reported - the process-yield cousin of kcat. | - |
Cofactor / Metals Ions | string | Required coenzymes and metal ions - the difference between a catalyst that works on paper and one that works in your buffer. About 40,907 activating-compound entries sit in this band. | - |
Organism / Source Tissue / Localization | string | Source taxonomy down to tissue and cellular compartment - the axis that turns one EC class into dozens of comparable entries. | Homo sapiens |
Sequence (UniProt Accession) / PDB ID | string | Cross-references into sequence and structure resources; more than 32 million linked sequence records across the resource. | - |
pH Optimum / Temperature Optimum / Stability | number | Operating windows: optimum pH, optimum temperature, plus solvent tolerance and storage stability - process-design parameters as fields. | - |
Application / Cloned / Engineering-Mutations / Disease | text | Applied annotations: documented uses, cloning records, crystallization states, engineered mutations and disease links. | - |
Additional fields on request | - | Pathway context, activators and inhibitors beyond the constants, molecular weight and subunit composition detail - pinned with examples once a sample names the families it needs. | - |
Coverage at a glance
| Dimension | Coverage |
|---|---|
| Geographic | Global - organisms across all kingdoms of life, from bacteria and archaea through plants, animals and human sources |
| Temporal | Continuously curated since 1987; numbered releases mark milestones (Release 2026.1 dated March 4, 2026); internal statistics series runs from 07.2007 to 01.2026 |
| Granularity | One entry per enzyme x organism x commentary x reference per field - measured values stay separated by source instead of averaged away |
Product specification
| Attribute | Value |
|---|---|
| Industry | Life & Health Insurance |
| Content types | Enzyme nomenclature and reaction equations, kinetic parameters (KM, kcat, kcat/KM, Ki, IC50, specific activity), cofactor and metal-ion requirements, organism-tissue-localization context, sequence and structure cross-references, pH/temperature/stability optima, applied annotations including mutations and disease links |
| Fields | Thirteen verified field families down to named substructures |
| Universe size | ~8,831 EC classes; ~40,894 KM values; ~94,073 turnover numbers; ~40,907 activating-compound entries; 32M+ linked sequence records |
| Source | BRENDA |
| Curator | Leibniz Institute DSMZ - an ELIXIR Core Data Resource and Global Core Biodata Resource |
| Quality score | 9/10 (catalog average 7.81 across 1,744 datasets) |
Questions buyers ask
What fields does the brenda enzyme database data include?
Thirteen verified field families: EC number with recommended, systematic and synonym names; reaction equations with substrates, products and CAS RN; KM values; turnover numbers; kcat/KM efficiency ratios; Ki and IC50 inhibition constants; specific activity; cofactors and metal ions; organism, tissue and localization; UniProt and PDB cross-references; pH, temperature and stability optima; and application, cloning, mutagenesis and disease annotations.
Does the data include both KM values and turnover numbers?
Yes - separately, not merged. Roughly 40,894 KM entries and some 94,073 turnover-number entries ship as distinct measured columns, each carrying its own commentary and cited reference, with kcat/KM efficiency ratios alongside so affinity and speed stay decomposable rather than collapsed into one summary figure.
Can entries be compared across organisms within one enzyme class?
That is the core capability. Entries exist at enzyme-organism-commentary-reference grain, so the same reaction measured across species sits side by side under one EC number. Ranking candidate hosts by catalytic efficiency becomes a sort on delivered rows instead of a cross-paper literature reconstruction.
Is enzyme inhibitor information included?
Yes. Ki values and IC50 concentrations arrive as numeric columns tied to named inhibitor compounds, with the experimental commentary attached. A sensitivity screen across a compound panel therefore runs as a filter on delivered records rather than a parsing project over published prose.
How current is the enzyme data?
Curation has run continuously since 1987, with numbered releases marking milestones: Release 2026.1, dated March 4, 2026, added 124 new and updated 750 enzyme classes. The resource's own statistics series reaches back to 07.2007, so coverage growth per class is itself chartable from the delivered records.
Can a sample be scoped to particular enzyme classes or organisms?
Yes - that is what the sample is for. Name EC classes, organisms, tissues or an applied angle such as engineered mutants, and real rows come back trimmed to that scope with the field dictionary pinned against the delivered records, so validation covers the exact extract a pipeline will receive.
Datasets that pair with this one
- BacDive — Bacterial Diversity Metadatabase (DSMZ) The descriptive half of the bench - 102,187 strains with husbandry and safety, zero catalysis numbers.
- ATCC — Biological Materials Catalog Once the screen picks a source organism, this is where authenticated materials come from.
- PRIDE Archive Proteomics Identifications Database Protein-level identifications from mass spectrometry - pairs with enzyme records when evidence moves from kinetics to peptides.
- Best life sciences tools & services datasets Where this record ranks among the eight scored - and what it beat.
- Life & health insurance data hub The full pooled industry view, from mortality tables to this biochemical outlier.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.