Datadory notebook
Raw Dna Sequencing Data: Dataset Structure and Field Coverage
Datadory delivers raw dna sequencing data covering comprehensive field definitions, entity mappings, and historical time series — structured for direct analytics and delivered on demand.
1,744 datasets. Pick your catch.
Which archives hold raw reads versus annotated sequences?
The split inside the International Nucleotide Sequence Database Collaboration is by processing stage, and picking the right layer saves days of reprocessing.
NCBI SRA stores original instrument output — reads, not consensus calls — nested four levels deep: Studies (SRP), Experiments (SRX), Samples (SRS) and Runs (SRR). A live esummary record shows the granularity concretely: run SRR40290107 inside experiment SRX34906359 of study SRP729053 carries 21,419,469 spots totalling 6,468,679,638 bases, sequenced on an Illumina NovaSeq 6000 with library strategy RNA-Seq in PAIRED layout. Scale is the headline: the largest public sequencing repository, holding tens of petabases across hundreds of millions of runs, with about 6 million samples linked to SRA runs from NCBI GEO alone. Every run also carries spot counts, base counts, byte size, instrument model and library strategy — WGS, RNA-Seq, WGA, Amplicon and ChIP-Seq among them.
The European Nucleotide Archive keeps that same raw-read layer plus assembled genomes and their annotation — tens of millions of run records organized per study, per sample, per experiment and per run, with INSDC heritage reaching back to the early 1980s. NCBI GenBank sits a step downstream: release 273.0 packages roughly 267 million annotated traditional records (~8.2 trillion bases) alongside about 6.4 billion WGS, TSA and TLS records (~51.8 trillion bases), each entry a self-describing flat file whose LOCUS line carries length, molecule type, topology, division code and modification date.
If you need expression matrices rather than reads, NCBI GEO is the functional-genomics layer above them: about 250,000 Series and 7.9 million Samples archived since 2000 under MIAME and MINSEQE standards, including 143,551 RNA-seq expression Series and 50,704 ChIP-seq binding Series.
When should you skip raw reads entirely?
CZ CELLxGENE Discover standardizes 2,216 datasets covering roughly 290 million cells into H5AD form re-processed to a versioned schema, with individual assets running up to about 18 GB and a hosted TileDB-SOMA Census sliceable by the Python or R package — cell-by-gene matrices instead of the FASTQ underneath, downloadable without credentials. Ensembl Genome Browser publishes whole-genome annotation per species as GTF, GFF3 and FASTA from its FTP tree on a roughly three-month cycle: release 116 arrived June 2026 carrying a human gene set of about 60,000 genes, so variant calling starts from a current reference rather than a fresh assembly job. Its REST API resolves stable identifiers directly — TP53 is ENSG00000141510.
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
NCBI Sequence Read Archive (SRA)
European Nucleotide Archive
last_updated
NCBI GenBank
LOCUS · ACCESSION · VERSION …+5 more
NCBI GEO
CZ CELLxGENE Discover
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sample