BacDive — Bacterial Diversity Metadatabase (DSMZ)

Datadory delivers bacdive bacterial diversity metadatabase dsmz data covering 102,187 standardized strain records - 22,126 of them type strains - with LPSN-backed taxonomy, NCBI tax ids, morphology, culture media with temperature and pH windows, isolation country and habitat, biosafety level, pathogenicity findings, 16S and genome accessions, and a citable per-strain DOI, delivered daily, weekly, or hourly.

What is the BacDive — Bacterial Diversity Metadatabase?

It is the reference card catalog of the microbial world, compiled by the Leibniz Institute DSMZ - an ELIXIR Core Data Resource and a Global Core Biodata Resource - and standardized to a degree few collections match: 102,187 bacterial and archaeal strain records, including 22,126 type strains, each carrying its own citable Digital Object Identifier such as 10.13145/bacdive1.20260601.11. Every record is organized into ten observed sections: general identification; name and taxonomic classification backed by LPSN down to species rank, synonyms included; morphology; culture and growth conditions; physiology and metabolism; isolation, sampling and environmental origin; interaction and safety; sequence information; genome-based predictions; and literature.

What separates this record from most of the shelf is that the standardization survives delivery. Colony color, shape and incubation period arrive as fields; culture conditions arrive as a named medium, a temperature range with optimum and a pH range; safety arrives as a biosafety level with comment sitting beside recorded pathogenicity toward human, animal and plant hosts. Datadory ships those ten sections as one typed table keyed on BacDive-ID.

The record pools into Datadory's life & health insurance data hub as one of the slice's 16 primary datasets. Get a sample of this dataset cut to the taxa, habitats or safety classes you actually screen.

What do BacDive sample rows look like?

One record per strain, straight from the August 2026 verification pass:

# delivered grain: one record per strain, sectioned by data category

BacDive-ID          : 132485
taxonomy            : Lysobacter tolerans
description         : aerobe, Gram-negative, rod-shaped bacterium that forms circular colonies
culture_medium      : LB (Luria-Bertani) MEDIUM
temperature_optimum : 28 C
isolation           : soil from hexachlorocyclohexane-contaminated dumpsite - India, Asia

# a second strain, same row shape
BacDive-ID      : 1      DSM-Number   : 2002
taxonomy        : Acetobacter aceti    NCBI tax id  : 435
biosafety_level : 1      doi          : 10.13145/bacdive1.20260601.11

Read the first row as three answers already decided for you: how the organism is classified (Lysobacter tolerans, aerobe, Gram-negative rods), how to grow it (LB medium, optimum 28 C), and where it came from (soil at a hexachlorocyclohexane-contaminated dumpsite in India, Asia). The second row shows the identifier spine every record rides: a DSMZ culture-collection number, an NCBI tax id, a biosafety level and a per-strain DOI on one line, so a citation, a join and a safety review all resolve without opening a browser tab.

What fields does the BacDive field dictionary define?

Twelve documented fields define the core record, each definition verified against live output during the August 2026 research pass - a bar met by 85.7 percent of the 1,744 datasets Datadory catalogs. They split into five jobs:

  • Identity: BacDive-ID, DSM-Number, the full scientific name with its LPSN-backed lineage and synonyms, the NCBI tax id and the per-strain DOI - the columns that make joins, citations and reclassification tracking possible.
  • Husbandry: recommended culture medium with whether growth occurs, plus temperature and pH ranges with their type.
  • Origin: isolation sample type with country, continent and latitude-longitude.
  • Handling risk: biosafety level with comment under the German risk-group classification, and recorded pathogenicity findings for human, animal and plant hosts.
  • Cross-reference: INSDC 16S rRNA and genome accession numbers with GC content.

The wider collection runs past any single page - fatty acid profiles, antibiotic susceptibility discs, metabolite utilization, spore formation and halophily all exist in the physiology section - so anything outside the twelve-column core ships as additional fields on request, documented in the same dictionary format once your sample names them.

How wide does coverage run, and at what grain?

Three chips summarize the footprint:

  • Geography: global - strains isolated worldwide, with country, continent and geolocation carried as fields rather than left in prose.
  • Temporal: a continuously curated collection; numbered releases carry milestone dates from December 2023 through December 2025, and the 2025 edition crossed the 100,000-strain mark while adding a genome browser.
  • Granularity: one record per strain (102,187), sectioned by data category.

Context for the collectors: this record scores 9/10 on Datadory's quality rubric against a catalogue mean of 7.81 - a tier shared by 534 of the 1,744 datasets catalogued - and it is one of the 16 primary datasets pooled in a slice averaging 8.31.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel and set the cadence to match the decision being fed; the rows arrive identical either way - cleaned, typed and keyed on BacDive-ID so the taxonomy spine, the culture panel and the safety panel line up without merge-key archaeology. Section structure resolves onto one stable record shape, multi-value physiology stays parseable, and every column is checked against the dictionary above before it reaches you. A sample scoped to your taxa, habitats or safety cuts ships first either way, with the field dictionary pinned against the delivered records.

Who uses this data, and for what?

Six jobs the strain ledger settles outright:

  1. Strain selection and safety triage - biosafety level and pathogenicity columns cut 102,187 organisms down to biosafety-level-1 candidates before anyone designs an experiment.
  2. Bioprospecting by habitat - isolation-source categories with geography turn 'which environments yield alkaliphiles' into a filter instead of a literature review.
  3. Culture-condition lookup - medium, temperature window and pH optimum for a hard-to-grow isolate, delivered as fields rather than footnote forensics.
  4. Taxonomic verification - LPSN-backed names with synonyms and NCBI tax ids keep a reclassification from breaking a pipeline mid-quarter.
  5. Feature engineering - morphology, growth ranges and habitat flags as model training columns; see ml model training.
  6. Citation-grade sourcing - the per-strain DOI gives a headline claim a stable identifier; see citation-grade research.

Market sizing on the tools and services that surround microbial work draws on the same collection; see market sizing.

Which personas get the most value?

Data Scientists & ML Engineers (relevance 3 of 3) get 102,187 typed strain records with morphology, growth ranges and habitat flags ready for feature work - see data scientists in life & health insurance. Developers & Data-Product Builders (3 of 3) stand up organism lookup and screening products on a stable BacDive-ID grain - see developers builders use cases. Journalists, Academics & Students (3 of 3) cite per-strain DOIs rather than someone's chart of a database - see journalists academics use cases. Market Researchers & Consultants (2 of 3) map microbial-diversity research activity by geography and habitat - see market researchers use cases.

The edge worth naming: this is descriptive metadata about organisms, not measured biochemistry. Kinetic values live in BRENDA Enzyme Database, and the two are compared directly in vs BRENDA Enzyme Database.

Which notes pair with this dataset?

Methodology note - records are community deposits curated at the Leibniz Institute DSMZ, so sections populate only where deposition data exist: one strain may carry fatty acid profiles while another stops at taxonomy and safety. Depth gets checked per record before anything ships rather than assumed uniform.

Completeness note - the twelve-column core travels on every record; section-level extras ship as additional fields on request, pinned with examples once a sample names the families it needs.

Provenance note - every record traces to a per-strain Digital Object Identifier, type-strain status where applicable, and INSDC sequence accessions - a chain of custody short enough to verify by eye.

Where to go next:

Browse the rest of the shelf on the life & health insurance data hub or the best life sciences tools & services datasets ranking, and read why a verified data dictionary matters downstream.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - twelve verified fields defining the strain record, with examples from the August 2026 capture
FieldTypeDefinitionExample
BacDive-IDintegerPrimary strain identifier and the join key across every delivered section.132485
DSM-NumberstringDSMZ culture-collection number where the strain is held there.2002
description / keywordstextAuto-generated strain summary sentence with keyword flags.Lysobacter tolerans UM1 is an aerobe, Gram-negative, rod-shaped bacterium...
Full scientific name + LPSN taxonomystringDomain-to-species ranks with synonyms, kept current against LPSN.Lysobacter tolerans — Pseudomonadota/Gammaproteobacteria
NCBI tax idintegerNCBI Taxonomy identifier with matching level, the crosswalk to sequence databases.435
Culture medium / growthstringRecommended medium name and whether growth occurs on it.LB (Luria-Bertani) MEDIUM — yes
Culture temp / pHstringTemperature and pH ranges with type (growth or optimum).25-40 growth; optimum 28 C; pH 3-10, alkaliphile
Isolation sample type / country / continent / lat-longgeoIsolation source description and geography with origin codes.soil from hexachlorocyclohexane-contaminated dumpsite — India, Asia
Biosafety level / risk groupintegerBiosafety level with comment under the German risk-group classification.1
Pathogenicity (human/animal/plant)stringRecorded pathogenicity findings per host type.None
16S / genome accessionsstringINSDC 16S rRNA and genome accession numbers with GC content.AF000162
doistringPer-strain Digital Object Identifier for citation and provenance.10.13145/bacdive1.20260601.11
Additional fields on request-Physiology-section extras - fatty acid profiles, antibiotic susceptibility discs, metabolite utilization, spore formation, halophily - pinned with examples once a sample names them.-

Coverage at a glance - geography, temporal span and granularity

DimensionCoverage
GeographicGlobal - strains isolated worldwide, with country, continent and geolocation carried as fields
TemporalContinuously curated; numbered releases with milestone dates from December 2023 through December 2025; the 2025 edition crossed 100,000 strains and added a genome browser
GranularityOne record per strain (102,187), sectioned by data category

Questions buyers ask

What does bacdive bacterial diversity metadatabase dsmz data contain?

Standardized records for 102,187 bacterial and archaeal strains, including 22,126 type strains, organized into ten sections: identification, LPSN-backed taxonomy with synonyms, morphology, culture and growth conditions, physiology and metabolism, isolation origin, interaction and safety, sequence information, genome-based predictions and literature. Each record carries a citable per-strain Digital Object Identifier.

Why do the 22,126 type strains matter?

Type strains are the nomenclatural reference points their species names hang on, which makes them the fixed anchors for any taxonomic verification work. With 22,126 of the 102,187 records holding type-strain status, a naming check resolves against the designated reference rather than a random deposited isolate.

Does the data include biosafety and pathogenicity information?

Yes. Every record carries a biosafety level with comment under the German risk-group classification, alongside recorded pathogenicity findings for human, animal and plant hosts. A screening pass can therefore cut 102,187 organisms to biosafety-level-1 candidates before any wet-lab planning begins.

Can the records be joined to sequence data?

They can. Records carry INSDC 16S rRNA and genome accession numbers with GC content, plus NCBI tax ids and LPSN-current scientific names with synonyms. Those cross-references let a strain table link into sequence repositories on stable identifiers instead of fuzzy name matching.

How current is the collection?

Curation runs continuously, with numbered releases marking milestones: dates run from December 2023 through December 2025, and the 2025 edition crossed the 100,000-strain mark while adding a genome browser. Per-record depth varies because sections populate only where deposition data exist.

Can a sample be scoped to particular taxa, habitats or safety classes?

Yes - that is what the sample is for. Name a genus, an isolation habitat such as soil or salterns, a geography or a biosafety cut, and real rows come back trimmed to that scope with the field dictionary pinned against the delivered records, so validation covers the exact extract a pipeline will receive.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing