NCBI GenBank

Datadory delivers biotechnology data covering NCBI GenBank, the NIH genetic sequence database: 267,383,895 annotated traditional records totaling about 8.2 trillion bases in release 273.0, plus roughly 6.4 billion WGS/TSA/TLS set-based records, each carrying its FEATURES annotation beside the sequence it describes. Delivered as API, files, or your warehouse.

What is NCBI GenBank data?

The annotated backbone of molecular biology. NCBI GenBank is the genetic sequence database run by the National Center for Biotechnology Information: the NIH's collection of publicly submitted DNA sequences, stored as self-describing records whose annotation travels with the sequence it describes. Together with DDBJ in Japan and the European Nucleotide Archive it forms the International Nucleotide Sequence Database Collaboration, so a submission taken in by any partner resolves here.

Scale first, because it decides architecture. Release 273.0, dated August 15, 2026, holds 267,383,895 traditional records totaling 8,236,878,868,450 bases, alongside roughly 6.4 billion set-based WGS, TSA and TLS records - whole-genome shotgun read sets, transcriptome shotgun assemblies and targeted-locus studies - covering another 51.8 trillion bases. Across both classes that is more than 6.6 billion sequences and 59 trillion bases, contributed continuously since 1982. Most records enter through direct researcher submissions and pass automated and manual quality processing before release. Get a sample of this dataset and inspect real records before committing pipeline time.

What does a sample of NCBI GenBank data look like?

One fully self-describing record, plus the release line that sizes the whole archive:

LOCUS       SCU49845    5028 bp    DNA    linear    PLN    29-OCT-2018
DEFINITION  Saccharomyces cerevisiae TCP1-beta gene, partial cds;
            and Axl2p (AXL2) and Rev7p (REV7) genes, complete cds.
ACCESSION   U49845
VERSION     U49845.1
SOURCE      Saccharomyces cerevisiae (brewer's yeast)

release inventory
release                 273.0
date                    August 15 2026
traditional_sequences   267383895
traditional_bases       8236878868450

Read the header line and the record explains itself. LOCUS states length (5,028 bp), molecule type (DNA), topology (linear), division (PLN - the plant-and-fungal grouping this yeast entry sits in) and modification date in a single stroke. SOURCE names the organism and drags its taxonomy lineage along behind it. VERSION pins the exact sequence so any citation of U49845.1 resolves to the same bases years later. The inventory underneath is the bulk consumer's view: a quarter-billion traditional records and eight-plus trillion bases sitting behind headers shaped exactly like this one. The example rows illustrate the record shape; field-level detail is documented below.

What fields does NCBI GenBank data include?

Nine core fields ride on every record, and they divide cleanly by job: identity (LOCUS, ACCESSION, VERSION), description (DEFINITION, KEYWORDS), provenance (SOURCE / ORGANISM, REFERENCE), annotation (FEATURES) and payload (ORIGIN). Definitions below follow the documented flat-file layout; the derived field families flagged in the footnote fold into your sample on request.

What does coverage look like across geography, time and granularity?

Geography - submissions arrive from laboratories worldwide and span every domain of life: bacterial, viral, primate, plant, fungal and synthetic constructs share one accession space, segmented by division codes rather than countries. Organism is the axis that matters, and SOURCE / ORGANISM ships the full taxonomy lineage on every record to slice it.

Temporal - a continuous archive since 1982, with each record carrying its own modification date on the LOCUS line and holdings versioned by numbered releases - 273.0 as of August 15, 2026. One compositional shift is worth planning around: since June 16, 2026 the archive no longer accepts personal sequence data from private individuals, so new human personal-genome entries route through institutional submitters.

Granularity - two tiers. Traditional records are per-sequence documents with per-feature annotation hanging off them, down to individual gene and CDS spans in FEATURES. Set-based records (WGS, TSA, TLS) hold their contigs as collections, which is why they dominate raw volume: about two dozen set-based entries exist for every traditional record. Scale check: 267,383,895 traditional records against 8,236,878,868,450 bases averages out near 31 kilobases per record - long enough that most analytical questions need the annotation index, not the full text.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Who uses this data, and for what?

  • Primer, probe and construct design - pull the annotated CDS spans for a gene via the /gene and /product qualifiers, take the flanks from FEATURES coordinates and the bases from ORIGIN, and VERSION pins the exact template the design was computed against.
  • Reproducibility audits - a published claim quoting a versioned accession resolves to the precise sequence its authors inspected; REFERENCE carries the AUTHORS, TITLE, JOURNAL and PUBMED trail back to the originating paper, so discrepancy becomes a diff rather than a correspondence chain.
  • Annotation-model training - FEATURES is supervised data at archive scale: millions of records with gene boundaries, coding regions and product labels aligned to the nucleotide text, ready-made inputs for gene-finders and function predictors.
  • Taxonomy and biodiversity rollups - SOURCE / ORGANISM lineages turn the archive into an organism-indexed census; division codes give a coarser but cheaper second cut.
  • Sequencing-activity intelligence - release-over-release record and base counts, split by division and technology class, price where deposition volume is flowing - a public proxy for where laboratories are spending instrument time.
  • Synthetic biology screening - division codes flag engineered constructs, and KEYWORDS plus DEFINITION make construct-level screening a filtered pull instead of a literature survey.

Which personas get the most value?

Bioinformaticians and computational biologists get records that need no side tables: sequence, coordinates and annotation arrive in one document, so a pipeline step is a parser rather than a join. Developers building genomics products get permanent, versioned accessions to key customer-facing features on - identifiers that survive reorganizations; see developers building in biotechnology. Data scientists and ML engineers get labeled spans at archive scale for training annotation and retrieval models; the data scientists working in biotechnology briefing goes deeper. Competitive-intel and diligence teams read release-over-release growth and division mix as deposition telemetry. Start from the biotechnology data hub or the best biotechnology datasets shortlist to see where this archive sits in the wider stack.

Which notes pair with this dataset?

Notes that pair well with this page:

  • Biotechnology data hub - the pooled industry view this record sits inside, alongside expression, structure and interaction corpora.
  • European Nucleotide Archive - the INSDC partner holding the same submission stream from the European side, with its own run-level grain.
  • The NCBI GenBank source profile - operator background, record classes and curation behavior of the underlying resource.
  • NCBI SRA read archives - where raw, unannotated read data for many of the same studies lives at run grain.
  • Ensembl genome browser data - reference annotation layered on top of these accessions for model organisms.
  • UniProt protein sequence data - the protein layer that CDS translations cross-reference, keyed by the same identifier discipline.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - nine core fields on every record; derived families fold below
fieldtypedefinitionexample
LOCUSstringRecord header line: locus name, sequence length, molecule type, topology, division code and modification date.SCU49845 5028 bp DNA linear PLN 29-OCT-2018
DEFINITIONtextSummary description of the record's sequence content.
ACCESSIONstringPrimary accession number, stable across versions.U49845
VERSIONstringAccession plus version number identifying the exact sequence.U49845.1
KEYWORDStextSubmitter-supplied keywords; often '.' when none.
SOURCE / ORGANISMstringSource organism name with its NCBI taxonomy lineage.Saccharomyces cerevisiae ... Eukaryota; Fungi; Ascomycota
REFERENCEtextLiterature citations with AUTHORS, TITLE, JOURNAL and PUBMED sub-fields.
FEATUREStextAnnotation table using INSDC feature keys (source, gene, mRNA, CDS) with location strings and qualifiers (/organism, /mol_type, /gene, /CDS, /product).
ORIGINtextThe nucleotide sequence itself, terminated by //.

Questions buyers ask

How big is NCBI GenBank?

Release 273.0, dated August 15, 2026, holds 267,383,895 traditional records totaling 8,236,878,868,450 bases, plus roughly 6.4 billion set-based WGS, TSA and TLS records covering another 51.8 trillion bases. Across both classes the archive exceeds 6.6 billion sequences and 59 trillion bases.

What are the WGS, TSA and TLS set-based records?

Whole Genome Shotgun read sets, Transcriptome Shotgun Assemblies and Targeted Locus Studies are held as sets of contigs rather than single flat-file records. They outnumber traditional records about two dozen to one, which means most sequence volume in the archive sits in the set-based tier.

How far back do the records go?

To 1982, when the collaboration began exchanging sequences, and contributions have continued without interruption since. A record's own LOCUS line dates it - the worked sample above was last modified October 29, 2018 - so era-slicing is a filter on a field rather than an inference from formatting.

How do GenBank versions work?

The ACCESSION stays stable forever while VERSION pins the exact sequence: U49845 identifies the record, U49845.1 identifies its first incarnation. Any reanalysis that quotes the versioned form resolves to the precise bases its authors inspected, which is what makes replication audits possible decades later.

How does GenBank relate to ENA and DDBJ?

They are the three partners of the International Nucleotide Sequence Database Collaboration: GenBank in the United States, ENA in Europe, DDBJ in Japan. Every partner shares each accepted submission with the others, so a sequence deposited anywhere in the partnership resolves in all three archives under shared accession namespaces.

What lives inside the FEATURES table?

Annotation keyed to the INSDC feature table: feature keys such as source, gene, mRNA and CDS, each with a location string and qualifiers like /organism, /mol_type, /gene, /CDS, /product and /db_xref. Gene boundaries, coding regions and product names therefore ride on the same record as the sequence they annotate.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing