Data source

Data from cBioPortal for Cancer Genomics, delivered clean.

5 datasets pulled from cBioPortal for Cancer Genomics's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.

  • 5 datasets
  • 1 industry
  • Real rows on request

What Datadory delivers from cBioPortal for Cancer Genomics

5
Biotechnology

cBioPortal for Cancer Genomics

Biotechnology

NCI Genomic Data Commons

Biotechnology Global submissions · Continuous archive from 2000 to present

NCBI GEO

Biotechnology Human populations worldwide · Publications from 2005 (the first genome-wide…

GWAS Catalog

Biotechnology Global - laboratory submissions worldwide via… · Current archive spans PXD001357 (submitted 20…

PRIDE Proteomics Archive data

Pick a catch, see the rows.

Name any cBioPortal for Cancer Genomics dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.

Get a sample

API, files, or your warehouse. Daily, weekly, or hourly.

Straight answers about cBioPortal for Cancer Genomics data

How big is the corpus Datadory delivers from cBioPortal?

As observed via our research pass in August 2026: 539 public cancer studies totaling 399,909 tumor samples across 897 cancer types, organized into 2,534 molecular profiles. All three counts climb as contributing institutions add studies, and each addition lands inside the same study-level dictionary.

Which cohorts sit inside the corpus?

Named contributors include the TCGA PanCancer Atlas cohorts, MSK-IMPACT clinical sequencing, AACR GENIE, PCAWG, CCLE and HTAN, alongside disease-specific institutional series. Study identifiers, citations, import dates and reference genomes are preserved per study, so cohort slices can respect institutional boundaries whenever the analysis demands it.

What data types does a single study carry?

Whatever that study collected, exposed as typed molecular profiles: mutations in MAF form, discrete and continuous copy number with segmentation, mRNA expression from RNA-seq or microarray, HM450 DNA methylation, CPTAC mass-spectrometry protein and phosphoprotein levels, RPPA panels, structural variants and mutational signatures. A flagship cohort may stack fifteen profiles; a small series may ship mutations and clinical data only.

Can the delivered data support survival analysis?

Yes. De-identified patients carry outcome attributes including overall-survival status and months, and those fields join onto the same barcode as the mutation and copy-number records. Testing whether an alteration associates with outcome becomes one query against one key instead of a merge across separate documents.