Data source

Data from CZ CELLxGENE (Chan Zuckerberg Initiative), delivered clean.

1 dataset pulled from CZ CELLxGENE (Chan Zuckerberg Initiative)'s releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.

  • 1 dataset
  • 1 industry
  • Real rows on request

What Datadory delivers from CZ CELLxGENE (Chan Zuckerberg Initiative)

1
Biotechnology Global published datasets spanning human · Published datasets from 2017 onward

CZ CELLxGENE Discover

Pick a catch, see the rows.

Name any CZ CELLxGENE (Chan Zuckerberg Initiative) dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.

Get a sample

API, files, or your warehouse. Daily, weekly, or hourly.

Straight answers about CZ CELLxGENE (Chan Zuckerberg Initiative) data

How big is the CZ CELLxGENE corpus?

Its curation inventory reported 2,216 datasets and roughly 289.5 million cells as of August 2026, while the portal homepage advertises 33M+ unique cells across 436 hosted collections. The two figures count different things - total cells including cross-study duplicates versus deduplicated unique cells - so pick one definition before quoting a scale number.

How far back does CZ CELLxGENE data go?

Published studies run from 2017 onward - effectively the entire modern single-cell RNA-seq era - and the catalog anchors to a long-term-support snapshot dated 2025-11-08, giving longitudinal work a fixed, citable vintage instead of a moving target.

Which organisms does CZ CELLxGENE cover?

Human and mouse dominate the corpus, and the 2025-11-08 long-term-support release added macaque (Macaca mulatta), common marmoset (Callithrix jacchus) and chimpanzee (Pan troglodytes). Organism is an NCBITaxon-tagged field on every record, so non-human cells are filtered explicitly rather than blended in unnoticed.

What can I join CZ CELLxGENE data against?

Ontology identifiers carry the joins: EFO for assays, Cell Ontology for cell types, MONDO for disease, UBERON for tissue, NCBITaxon for organism, HANCESTRO for self-reported ethnicity, and ENSEMBL gene names at the feature level. Those keys resolve deterministically into pathway databases, GWAS catalogs and protein resources without fuzzy string matching.