Biotechnology Data Provider: 24 Cataloged Datasets · Head-to-head
NCBI GenBank vs NCI Genomic Data Commons
Which biotechnology data provider: 24 cataloged datasets data fits your job: NCBI GenBank, or NCI Genomic Data Commons. API, files, or your warehouse. Daily, weekly, or hourly.
NCBI GenBank
NCI Genomic Data Commons
Where the fields line up
No shared field names. These two answer different questions.
| Field | NCBI GenBank | NCI Genomic Data Commons |
|---|---|---|
LOCUS | documented | not in this set |
DEFINITION | documented | not in this set |
ACCESSION | documented | not in this set |
VERSION | documented | not in this set |
KEYWORDS | documented | not in this set |
SOURCE / ORGANISM | documented | not in this set |
REFERENCE | documented | not in this set |
FEATURES | documented | not in this set |
ORIGIN | documented | not in this set |
case_id | not in this set | documented |
submitter_id | not in this set | documented |
project.project_id | not in this set | documented |
Coverage, side by side
| NCBI GenBank | NCI Genomic Data Commons | |
|---|---|---|
| Geographic | Global submissions from laboratories worldwide, spanning every domain of life | United States-centric cancer cohorts with international contributions mixed in |
| Temporal | Continuous since 1982; bimonthly numbered releases (current Release 273.0, August 15, 2026) with daily incremental updates | TCGA samples collected 2005-2013 with later programs ongoing; commons launched 2016, currently Data Release 46.0 (August 10, 2026) |
| Granularity | One self-describing record per sequence with per-feature annotation hanging off it; set-based WGS/TSA/TLS collections alongside | One row per case, with per-file children across genomic, transcriptomic, epigenomic, proteomic, imaging and clinical data types |
What each contains
Pick by fit, not by loyalty.
| NCBI GenBank | NCI Genomic Data Commons | |
|---|---|---|
| Publisher | National Center for Biotechnology Information (NCBI), part of the National Institutes of Health - archiving community submissions since 1982 | National Cancer Institute (NCI) research program - the unified cancer genomics repository launched in 2016 |
| Subject lens | Annotated DNA sequence records for any organism anyone has submitted, with feature-table annotation traveling inside every record | Harmonized cancer genomics: sequencing alignments, variant calls, expression, methylation and pathology slides anchored to clinical and biospecimen data per patient |
| Unit of analysis | One self-describing record per sequence with per-feature annotation hanging off it; set-based WGS/TSA/TLS collections alongside | One row per case, with per-file children across genomic, transcriptomic, epigenomic, proteomic, imaging and clinical data types |
| Geographic coverage | Global submissions from laboratories worldwide, spanning every domain of life | United States-centric cancer cohorts with international contributions mixed in |
| Temporal coverage | Continuous since 1982; bimonthly numbered releases (current Release 273.0, August 15, 2026) with daily incremental updates | TCGA samples collected 2005-2013 with later programs ongoing; commons launched 2016, currently Data Release 46.0 (August 10, 2026) |
| Formats | GenBank flat file, ASN.1, FASTA | BAM, VCF, MAF, TSV, JSON, IDAT, SVS pathology slides, BCR XML |
| Scale | 267,383,895 traditional records (8,236,878,868,450 bases) plus roughly 6.4 billion set-based records covering another 51.8 trillion bases | 50,571 cases and 1,337,360 files, led by VCF (about 364,000), BAM (about 200,000) and TSV (about 199,000) |
| Best for | Primer and construct design, similarity search, taxonomy rollups, annotation-model training, reproducibility audits | Biomarker discovery, cohort construction, survival modeling, pathology-imaging AI, translational research |
What each does better
NCBI GenBank
Scale of raw sequence nothing else in the pairing approaches. Release 273.0 holds 267,383,895 traditional records totaling 8,236,878,868,450 bases, plus roughly 6.4 billion set-based WGS, TSA and TLS records covering another 51.8 trillion bases - more than 6.6 billion sequences and 59 trillion bases overall. The GDC's entire holdings are 1,337,360 files. Any question whose unit is a molecule, not a person, has only one answer here.
Every domain of life, continuously since 1982. Submissions arrive from laboratories worldwide across bacterial, viral, plant, fungal, primate and synthetic divisions - SOURCE / ORGANISM ships the full lineage on every record, so the archive slices by organism rather than by country. The GDC covers one disease family, however deeply.
Annotation travels with the sequence it describes. FEATURES keys gene, mRNA and CDS spans to the INSDC feature table with /gene, /product and /mol_type qualifiers, so primer design, construct sourcing and reproducibility audits resolve inside a single document - no join required. See sequence annotation for how this record structure gets used.
NCI Genomic Data Commons
Clinical context rides on every molecular record. Each case carries primary_site ('Bronchus and lung'), disease_type ('Adenomas and Adenocarcinomas'), diagnoses.primary_diagnosis ('Adenocarcinoma, NOS'), diagnoses.ajcc_pathologic_stage ('Stage IA'), demographic.race, demographic.vital_status and days_to_lost_to_followup - the variables biomarker discovery joins against, arriving pre-linked instead of assembled from separate papers. GenBank has no patient dimension whatsoever.
One harmonized pipeline instead of thousands of submission habits. Alignments land as BAM, variant calls as VCF and MAF, expression quantification standardized, methylation as IDAT, whole-slide pathology images as SVS, clinical data as TSV, JSON or XML - all reprocessed against a common pipeline and described by one versioned data dictionary. GenBank deliberately preserves each submitter's own annotation choices.
Cohort construction is a filter, not a build. Fifty thousand seven hundred sixty-one cases across named programs - TCGA with 11 projects, TARGET with 6, CPTAC-2 and CPTAC-3 proteogenomics, plus CGCI, BEATAML, Exceptional Responders, MATCH, ALCHEMIST and more - let a study size a recruitable cohort by project code, site, disease type and stage before any protocol commits.
Where they're equivalent
More than their different shapes suggest. Both score 10/10 on Datadory's rubric against the catalog's 7.81 average, and both field dictionaries were verified during research - a bar 85.7 percent of the catalog clears. Both publish numbered, datable releases you can pin a pipeline to for reproducibility: GenBank 273.0 (August 15, 2026) and GDC Data Release 46.0 (August 10, 2026). Both are United States government-run repositories operating at archive scale, both grow by continuous contribution rather than periodic re-issue, and both ship documented machine-readable structures rather than narrative pages alone - flat-file and ASN.1 records on one side, TSV, JSON and XML entities on the other.
The verdict
Verdict: sample both, pick by fit - they are different instruments pointed at the same industry.
Take NCBI GenBank if your question names a molecule. Primer, probe and construct design off curated CDS spans, similarity-search pipelines, taxonomy and biodiversity rollups, annotation-model training at archive scale, synthetic-biology part sourcing - anything answered by versioned accessions with the annotation attached. Accept molecule-first design: no outcome, stage or treatment detail travels with any record.
Take NCI Genomic Data Commons if your question names a patient. Biomarker candidates checked for prognostic signal, cohort feasibility by stage and site, survival modeling on outcome-labeled cases, pathology-imaging AI paired with coded ground truth - anything needing molecular evidence joined to clinical reality. Accept the frame: one disease family, human, harmonized to someone else's pipeline.
Data scientists working in biotechnology usually reach for the GDC first on supervised problems and GenBank for training corpora; developers building genomics products tend toward GenBank for reference features and permanent identifiers, then add the commons when the product needs an outcome attached.
Sample both, pick by fit. See NCBI GenBank · See NCI Genomic Data Commons
Or take both in one feed
Yes - they stack because they occupy different layers of the same system. A defensible workflow: let GenBank supply the annotated reference world - resolve gene and CDS coordinates from FEATURES entries, with VERSION pinning the exact template - then read the GDC underneath it for the observations layered on the human genome, mapping each variant call onto the annotated transcript to name the affected product. Reverse the flow for target discovery: pick candidate sequences in GenBank first, then check which cases in the commons carry variation there and what happened to those patients.
Two alignments decide whether the merge holds. First, identifiers: GenBank speaks accession.version while the GDC hands over universally unique case identifiers and program-native barcodes, so a crosswalk is mandatory - the join runs through gene and transcript coordinates, not through either repository's native key. Second, units of analysis differ in kind, so keep the layers separated until analysis time rather than summing across cases and sequences as if they were rows of one table. Browse the rest of the shelf at the biotechnology data hub.
Datadory ships either record alone or both merged onto one calendar, delivered daily, weekly, or hourly - your call. Or take both in one feed.
API, files, or your warehouse. Daily, weekly, or hourly.
Fair questions
Is NCBI GenBank better than NCI Genomic Data Commons?
Better at different jobs. GenBank owns the molecule: more than 6.6 billion annotated sequences - 267,383,895 traditional records and 59 trillion bases total - from every domain of life since 1982, each carrying its own FEATURES annotation. The GDC owns the patient: 50,571 harmonized cancer cases with diagnosis, AJCC stage and vital status joined to 1.3 million-plus files. Sample both, pick by fit.
Do the two datasets cover the same ground?
Only partly. The overlap is human sequence information, and even there they disagree by design: GenBank stores the annotated sequence document exactly as submitted, while the GDC reprocesses tumor-derived molecular files through one harmonized pipeline and anchors them to clinical outcomes. Outside that overlap they diverge completely - multi-kingdom sequence archive on one side, cancer cohort knowledge base on the other.
Which dataset has broader organism coverage?
NCBI GenBank, by an enormous margin: submissions span bacterial, viral, plant, fungal, primate and synthetic divisions - every domain of life - each record carrying its full lineage in SOURCE / ORGANISM. The NCI Genomic Data Commons covers Homo sapiens only, focused on malignant disease across its contributing programs.
Which should anchor a biomarker discovery screen?
Start with the GDC. Stage, vital status, follow-up duration and coded diagnoses turn 50,571 cases into a filterable cohort, and harmonized variant calls arrive joined to those outcomes. Then take candidate genes into GenBank to pull annotated transcript structure - CDS spans, products, accession versions - for the regions your signal points at. Reversing the order means reading sequence documents before any clinical hypothesis exists.
Can Datadory deliver both datasets together?
Yes - alone or merged onto one calendar, delivered daily, weekly, or hourly, your call. Each arrives normalized to its documented field dictionary (nine fields on the GenBank side, thirteen on the GDC side) with sample rows for inspection before anything ships; the joining work is transcript-level identifier alignment between the two, handled in the merge. Or take both in one feed.