Biotechnology Data Provider: 24 Cataloged Datasets · Head-to-head

OpenFDA vs cBioPortal

Which biotechnology data provider: 24 cataloged datasets data fits your job: OpenFDA, or cBioPortal. API, files, or your warehouse. Daily, weekly, or hourly.

Biotechnology Data Provider: 24 Cataloged Datasets Primarily United States jurisdiction · Drug adverse events reach back to 2004 with quarterly partitions from 2004 Q3

OpenFDA

Biotechnology Data Provider: 24 Cataloged Datasets Global contributor institutions - MSK · Studies imported continuously

cBioPortal

Where the fields line up

No shared field names. These two answer different questions.

Field OpenFDA cBioPortal
safetyreportid documented not in this set
receivedate documented not in this set
seriousness_criteria documented not in this set
patient_reactions documented not in this set
patient_drugs documented not in this set
occurcountry documented not in this set
reporter_country documented not in this set
label_set_id documented not in this set
effective_time documented not in this set
boxed_warning documented not in this set
indications_and_usage documented not in this set
active_ingredient documented not in this set

Coverage, side by side

OpenFDA cBioPortal
Geographic Primarily United States jurisdiction, with foreign event occurrences captured via occurcountry Global contributor institutions - MSK, Broad, Dana-Farber, TCGA network, AACR GENIE consortium - restricted to human cancer
Temporal Drug adverse events reach back to 2004 with quarterly partitions from 2004 Q3; labels and listings span current plus historical holdings Studies imported continuously, portal-wide re-import January 2026, additions through August 2026; contents date to original publications

What each contains

Pick by fit, not by loyalty.

OpenFDA cBioPortal
Publisher The U.S. Food and Drug Administration, through its openFDA platform A cancer genomics portal begun at Memorial Sloan Kettering, carried forward by a global research community
Subject lens Post-market regulation: adverse event reports, structured product labels, NDC listings, enforcement actions across drugs, devices, food, cosmetics, animal health and tobacco Cancer genomics research: mutations, copy number, expression, methylation, protein quantification and clinical outcomes across 897 observed cancer types
Geographic coverage Primarily United States jurisdiction, with foreign event occurrences captured via occurcountry Global contributor institutions - MSK, Broad, Dana-Farber, TCGA network, AACR GENIE consortium - restricted to human cancer
Temporal coverage Drug adverse events reach back to 2004 with quarterly partitions from 2004 Q3; labels and listings span current plus historical holdings Studies imported continuously, portal-wide re-import January 2026, additions through August 2026; contents date to original publications
Detail level One record per regulatory object: event report, label version, product listing, recall action Per-sample, per-gene alteration records nested within 539 studies, plus per-patient clinical attributes and timelines
Formats JSON responses plus whole-study files in TSV, MAF and SEG layouts
Field dictionary 25 documented fields, verified during research 11 documented fields, verified during research

What each does better

OpenFDA

Scale measured in millions of reports. Roughly 55 million records across 29 datasets: device adverse events at about 25.7 million, drug adverse events at about 20.7 million, unique device identifier records near 5.1 million, and about 262,000 versions of structured product labels. Nothing in the cancer genomics world approaches that volume of individual real-world reports.

Breadth beyond oncology entirely. Six territories - human drugs, medical devices, food, cosmetics, animal and veterinary events, and tobacco research - plus reference data such as the NDC directory and Drugs@FDA approvals. A safety or competitive question about any regulated product lands here first; see FAERS adverse event report.

The regulator's own text. Label versions carry the authoritative sections verbatim - boxed_warning, indications_and_usage, active_ingredient - alongside enforcement and recall actions. When the question is what the agency actually says about a product, this is the primary record.

cBioPortal

Molecular depth under one schema. Seven alteration families - mutations (MAF), discrete and continuous copy number, segmentation, mRNA expression, DNA methylation, protein and phosphoprotein quantification, structural variants - plus mutational signatures, treatments and timeline events, all joined to de-identified clinical attributes. One query model covers 2,534 molecular profiles.

Cohorts that define the field. TCGA PanCancer Atlas, MSK-IMPACT clinical sequencing, AACR GENIE, PCAWG, CCLE and HTAN sit alongside disease-specific institutional series, spanning 897 observed cancer types and 399,909 samples.

OpenFDA records what was reported about a product; cBioPortal records what happened inside the tumor, sample by sample, down to read counts supporting each variant allele.

Where they're equivalent

More than their subjects suggest. Both field dictionaries were verified during research, and both sit at the top of their slice - 10/10 and 9/10 against a biotechnology average of 8.92 and a catalog average of 7.81. Both attribute every row to an accountable party, whether the reporter country and qualification on an event report and the labeler behind an NDC listing, or the contributing institution and citation behind a study. Both organize their rows under controlled vocabularies - MedDRA terms and marketing categories here, OncoTree cancer types and MAF variant classes there. And both pair hard structure with narrative text: label sections and study descriptions ride along beside the tabular fields rather than living in separate documents.

The verdict

Verdict: sample both, pick by fit - they are different instruments pointed at the same industry.

Take OpenFDA if your question contains a product. How many serious adverse events mention this therapy, which companies list an equivalent generic, what changed between two versions of a label, where recalls cluster - anything answered by regulatory records accumulated after market entry. Accept that reports are spontaneous submissions shaped by reporting behavior, not incidence rates.

Take cBioPortal if your question contains a tumor or a gene. How often is this alteration present in a cohort, do these two genes co-occur, do patients carrying it survive longer - anything answered by molecular profiling joined to outcomes. Accept that depth stops at the study boundary: 539 curated universes rather than one continuous register.

Data scientists usually start with the genomics side; competitive intel and product teams usually start with the regulator's.

Sample both, pick by fit. See OpenFDA · See cBioPortal

Or take both in one feed

Yes - they stack because they occupy different layers of the same evidence chain, molecule to market. A defensible workflow: define the biology in cBioPortal first - which alterations mark a responsive population, which cohorts carry them, what survival looks like - then read OpenFDA underneath for the marketed-product reality: the label text and warning history for therapies aimed at those targets, their adverse event profile, and competing NDC listings sharing the same active ingredient.

Two alignments decide whether the merge holds. First, identifiers: there is no shared row key, so a crosswalk through drug-to-target mappings is mandatory before an alteration can meet a label. Second, tempo: event reports accumulate on a filing clock while genomic studies arrive as curated imports, so any joined artifact should record which vintage of each it used. Neither record needs the other to be useful, but the pair covers biology and regulation of the same medicines in a way neither half manages alone. Browse the rest of the shelf at the biotechnology data hub.

Datadory ships either record alone or both merged onto one calendar, delivered daily, weekly, or hourly - your call, normalized to the verified field dictionaries above, with sample rows for inspection before anything ships. Or take both in one feed.

API, files, or your warehouse. Daily, weekly, or hourly.

Fair questions

Is cBioPortal better than OpenFDA?

Better at different jobs. OpenFDA owns regulation: about 55 million FDA records spanning adverse events, product labels, NDC listings and recalls across drugs, devices and food. cBioPortal owns tumor biology: 539 public studies over 399,909 samples of mutations, copy number and outcomes. One records what was reported about products; the other records what happened inside cancers.

Do OpenFDA and cBioPortal cover the same ground?

Only at the edges. Both verify their field dictionaries, score at the top of the biotechnology slice (10/10 and 9/10), attribute rows to accountable parties, and encode rows in controlled vocabularies. But their units never meet: OpenFDA's rows are regulatory objects - reports, labels, listings - while cBioPortal's rows are alterations observed in de-identified tumor samples.

Which dataset is bigger?

OpenFDA, by raw volume: roughly 55 million records across 29 datasets, led by 25.7 million device adverse events and 20.7 million drug adverse events. cBioPortal counts 399,909 samples across 539 studies and 2,534 molecular profiles - smaller in rows, far denser per row, with individual expression matrices running to multiple gigabytes.

Which should a pharmacovigilance team sample first?

Start with OpenFDA - its drug and device adverse event collections carry coded reactions, seriousness criteria including death, reporter qualifications and occurrence countries, which is precisely the raw material of signal detection. Add cBioPortal second when the team needs the molecular substrate: which tumor alterations a therapy targets and how common those alterations are by cancer type.

Which should a biomarker discovery team sample first?

Start with cBioPortal. Bring in OpenFDA second, to see how the marketed therapies against those targets are labeled and what event history follows them.

Can Datadory deliver both datasets together?

Yes - alone or merged onto one feed, delivered daily, weekly, or hourly, your call. Each arrives normalized to its verified field dictionary (25 documented fields on the OpenFDA side, 11 on cBioPortal) with sample rows for inspection first. The only joining work is the drug-to-target crosswalk between them, handled in the merge. Or take both in one feed.