Data source

Data from NCBI Sequence Read Archive, delivered clean.

4 datasets pulled from NCBI Sequence Read Archive's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.

  • 4 datasets
  • 1 industry
  • Real rows on request

What Datadory delivers from NCBI Sequence Read Archive

4
Biotechnology Global - submissions from sequencing centers · Continuous archive since roughly 2008 through…

NCBI Sequence Read Archive (SRA)

Biotechnology Global submissions · Continuous archive from 2000 to present

NCBI GEO

Biotechnology Global submissions spanning all domains of li… · Continuous archive since 1982

NCBI GenBank

Biotechnology Global - submissions from international seque… · Holdings extend from the early 1980s (INSDC h…

European Nucleotide Archive

Pick a catch, see the rows.

Name any NCBI Sequence Read Archive dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.

Get a sample

API, files, or your warehouse. Daily, weekly, or hourly.

Straight answers about NCBI Sequence Read Archive data

How big is the NCBI Sequence Read Archive?

The largest public store of raw sequencing data in existence: tens of petabases spread across hundreds of millions of runs, accumulated continuously since roughly 2008. About 6 million samples link into these runs from NCBI GEO alone, which gives a sense of how much downstream work sits on top of the raw layer.

What is the difference between a study, an experiment, a sample and a run?

They are the four levels of the archive's hierarchy. A study (`SRP`) is the parent project, an experiment (`SRX`) is one sequencing setup inside it, a sample (`SRS`) is the biological material, and a run (`SRR`) is one execution of the machine producing reads. Datadory ships the whole spine, so any level resolves by join.

Does the archive hold raw reads or processed results?

Raw reads - the unprocessed signal straight off the instruments, plus alignment information. Processed expression matrices live in GEO and annotated consensus sequences in GenBank; both sit one layer above this archive, and both trace back to runs here when a claim needs checking at the source.

What can I join this data against?

Taxonomy identifiers and scientific names match into any organism-keyed collection; BioProject and BioSample cross-references reach the rest of the NCBI ecosystem; GEO sample links connect raw reads to their published matrices. Any system keyed on run accessions converges here with a single hash join.