Data source

Data from NCBI GenBank, delivered clean.

6 datasets pulled from NCBI GenBank's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.

  • 6 datasets
  • 1 industry
  • Real rows on request

What Datadory delivers from NCBI GenBank

6
Biotechnology Global submissions spanning all domains of li… · Continuous archive since 1982

NCBI GenBank

Biotechnology Global - submissions from international seque… · Holdings extend from the early 1980s (INSDC h…

European Nucleotide Archive

Biotechnology Global - submissions from sequencing centers · Continuous archive since roughly 2008 through…

NCBI Sequence Read Archive (SRA)

Biotechnology Global submissions · Continuous archive from 2000 to present

NCBI GEO

Biotechnology Not applicable (reference genomes) · Versioned releases

Ensembl Genome Browser

Biotechnology

UniProt

Pick a catch, see the rows.

Name any NCBI GenBank dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.

Get a sample

API, files, or your warehouse. Daily, weekly, or hourly.

Straight answers about NCBI GenBank data

How many datasets does Datadory deliver from NCBI GenBank?

One catalogued collection in Biotechnology, and it is the archive's full holding: roughly 267,383,895 traditional records and about 6.4 billion set-based WGS, TSA and TLS records as of release 273.0. It ships with a verified field dictionary, worked examples and sample rows before you commit.

How big is NCBI GenBank?

Release 273.0, dated August 15, 2026, holds 267,383,895 traditional records totaling 8,236,878,868,450 bases, plus roughly 6.4 billion set-based records covering another 51.8 trillion bases. Across both classes the archive exceeds 6.6 billion sequences and 59 trillion bases.

How far back do the records go?

To 1982, continuously, through the present. Each record carries its own modification date on the LOCUS line - the worked example above was last modified October 29, 2018 - so era-aware filtering is a field filter rather than guesswork.

How do GenBank versions work?

The ACCESSION stays stable forever while VERSION pins the exact sequence: U49845 identifies the record, U49845.1 identifies its first incarnation. Any reanalysis quoting the versioned form resolves to the precise bases its authors inspected, which is what makes replication audits possible decades later.