The Cancer Imaging Archive (TCIA)

Datadory delivers the cancer imaging archive tcia data covering roughly 240 de-identified collections of DICOM radiology and digital pathology, from named cohorts such as TCGA-LUAD and CBIS-DDSM up to 26,254 subjects in NLST, each joined to outcomes, genomics and expert segmentations. Delivered daily, weekly, or hourly.

Where it covers
Global contributor institutions; primarily US NCI-sponsored trials and international research cohorts
How far back
Collections added continuously since 2010; per-row Updated dates, 26 of 240 collections still Ongoing; acquisition windows vary by trial cohort
How fine
Patient, then imaging study, then DICOM series, then image instance, with per-collection clinical and genomic annotation tables

What is The Cancer Imaging Archive (TCIA)?

The pixels behind most published cancer-AI results, opened as a structured catalog instead of a download folder. The Cancer Imaging Archive (TCIA) is the National Cancer Institute's de-identified archive of medical images of cancer, funded through its Cancer Imaging Program and hosted at the University of Arkansas for Medical Sciences since October 2015. It carries ISSN 2474-4638, which is unusual for a dataset and tells you how deliberately it is curated.

Scale first: the Browse Collections table lists 240 rows - 211 marked Complete, 26 Ongoing - and each row is a patient cohort united by disease (lung cancer), modality (MRI, CT, digital histopathology) or research focus. Cohort sizes run from dozens of subjects up to 26,254 in NLST. What separates this record from every registry in the industry is what hangs off the pixels: where a cohort supplies them, outcome tables, treatment details, genomics and expert segmentations join directly to the imagery, so a model trains on pairs rather than on an annotation slog.

It scores 9 out of 10 on Datadory's quality rubric - against a catalog-wide average of 7.81 - with field definitions verified down to individual DICOM tags during research. It is one of six datasets tagged directly to health care technology, and it ranks 7th on our best health care technology datasets list.

What do sample rows look like?

One row per collection from the catalog's Browse Collections table, exactly as it lands:

# collection row - cohort inventory as shipped
collection   : PSMA-PET-CT-Lesions
cancer_type  : Prostate Cancer          species : Human
locations    : Prostate
subjects     : 378
data_types   : Demographic, Protocol, CT, PT, SEG
status       : Complete

collection   : BraTS-PEDs
cancer_type  : Diffuse Midline Glioma   locations : Brain
subjects     : 457                      data_types: MR, Segmentation, Other, Demographic

collection   : HANCOCK
cancer_type  : Head and Neck Squamous Cell Carcinoma
locations    : Oropharynx, Hypopharynx, Larynx, Oral Cavity
subjects     : 763
data_types   : Whole Slide Image, Tissue Microarray

collection   : UTSW-Glioma
cancer_type  : Glioma                   locations : Brain
subjects     : 625
data_types   : Demographic, Diagnosis, Treatment, Molecular Test, MR, Segmentation

# scale anchor
NLST         : subjects 26,254 - largest subject count in the table

Read it as two layers, not one. The first layer is the cohort inventory: subjects sizes the cohort, data_types names everything attached (imaging modalities like CT, PT and MR, but also Demographic, Diagnosis, Treatment and Molecular Test tables), and status says whether the collection has stopped growing. The second layer sits underneath: every collection decomposes into patients, then imaging studies, then DICOM series, then image instances - so PSMA-PET-CT-Lesions' 378 subjects fan out into CT and PET series plus expert SEG objects keyed to the same patient IDs. HANCOCK shows the other shape this record takes: whole slide images and tissue microarrays instead of radiology. Same schema, different tissue.

What fields does the dataset include?

Ten fields anchor the dictionary, verified against the archive's own documentation during research rather than inferred from column headers. They split cleanly in half: five describe the cohort you are joining to, and five describe the DICOM hierarchy underneath it. Per-collection annotation tables - outcomes, treatment details, genomics, expert segmentations - ship alongside the imagery where the cohort supplies them, documented field by field when you request a sample.

What does coverage look like across geography, time and granularity?

Geography - global contributor institutions, weighted toward US NCI-sponsored trials and international research cohorts. This is not population surveillance: cohorts self-select into the archive through research programs, so counts measure contribution, never incidence.

Temporal - collections have been added continuously since 2010, and acquisition dates vary by cohort because they belong to trials that ran years apart. Every Browse Collections row is stamped with an Updated date beside its Status, and 26 collections are still marked Ongoing - growing archives, not frozen ones. Match the window to the question before matching anything else.

Granularity - four levels deep: patient, then imaging study, then DICOM series, then image instance, with per-collection clinical and genomic tables hanging off the same keys. That depth is why computer-vision work starts here while country-year policy series live elsewhere in the pool; a segmentation model needs instance-level pixels, and nothing shallower can fake it.

Set against the wider catalog: average quality across all 1,744 datasets Datadory tracks is 7.81, and only 89 of them score at or above this record's tier of documentation depth.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Hourly suits an active research program tracking a collection that just flipped to Ongoing-to-Complete transitions worth catching the week they happen. Daily suits a model-training pipeline that wants new cohorts and new annotation tables as they land. Weekly suits literature-review cadence, where the point is keeping a cohort inventory current without babysitting it. However you cut it, the join keys stay stable: PatientId inside a collection, SeriesInstanceUID down at the pixel level, collection everywhere above - snapshots concatenate into history without remapping.

Who uses this data, and for what?

  • Training cancer-detection models - radiology cohorts arrive paired with expert segmentations and outcome labels, so the annotation work that usually eats months is already done; see the ML model training use case.
  • Validating segmentation algorithms - expert-labeled series across CT, MR and PET give benchmark sets with ground truth attached, the difference between publishing a number and defending one.
  • Radiomics and multimodal studies - linking image features to genomics and treatment outcomes inside one keyed schema turns 'does the scan predict the mutation' into a query instead of a collaboration request.
  • Digital pathology pipelines - HANCOCK-style whole slide images and tissue microarrays feed histology classifiers without a microscope stage in between.
  • Teaching files and academic publication - de-identified, DOI-cited collections give educators and journal reviewers a stable referenceable corpus; see the citation-grade research use case.

Which personas get the most value?

Data scientists and ML engineers get the headline: labeled imaging at cohort scale, with the label tables already keyed to the patient IDs. Developers building data products get a typed, documented schema down to DICOM tags - the difference between an integration and a science project. Journalists, academics and students publish against an NCI-funded archive whose provenance survives peer review. Market researchers and consultants scope oncology AI markets from a cohort inventory with real subject counts instead of press-release numbers. Investors and quant researchers read which institutions contribute which cohorts when diligence on an imaging startup claims a proprietary moat that is actually a public one. Persona workflows live at data scientists x health care technology, developers & builders x health care technology and competitive intel & product teams x health care technology.

What should I know before requesting a sample?

Three things worth knowing upfront. First, coverage is cohort-shaped, not census-shaped: contributors self-select through research programs, so the archive measures what has been contributed, never disease frequency - do not read subject counts as prevalence. Second, the archive publishes per-collection subject counts but no archive-wide image total, so any 'millions of images' figure you see quoted elsewhere is arithmetic someone else did; we size samples per collection instead. Third, DUA-required cohorts moved to dbGaP-managed access after April 2025 NIH policy changes, so open and restricted populations are distinct - name the cohorts you want and the sample comes back scoped honestly. None of these are defects we smooth over; they are the terrain, and the sample shows them to you unchanged.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - ten verified fields spanning the cohort inventory and the DICOM hierarchy beneath it
FieldTypeDefinitionExample
collectionstringName of the patient cohort - the top-level key every other field hangs off.TCGA-LUAD
PatientIdstringDe-identified patient identifier (DICOM tag 0010,0020) that keys patients within a collection.Calc-Test_P_00038_LEFT_CC
StudyInstanceUIDstringUnique identifier for one imaging study (DICOM tag 0020,000D); the middle layer between patient and series.<returned in your sample>
SeriesInstanceUIDstringUnique identifier for an imaging series (DICOM tag 0020,000E); the key image instances resolve against.1.3.6.1.4.1.14519.5.2.1.2135...
ModalityenumImaging modality of the series (DICOM tag 0008,0060) - the field that splits CT from MR from PET from mammography.MG
BodyPartExaminedstringAnatomical region imaged (DICOM tag 0018,0015); pairs with Modality to define a cohort's clinical surface.BREAST
ManufacturerstringMaker of the imaging equipment (DICOM tag 0008,0070); the field scanner-heterogeneity analyses group by.<returned in your sample>
ImageCountintegerNumber of images in the series - how instance-level volume reads off a single row.<returned in your sample>
SubjectsintegerSubject count for the collection, from the Browse Collections table; the cohort-sizing column.26254

Questions buyers ask

How many collections are in The Cancer Imaging Archive?

The Browse Collections table lists 240 rows - 211 marked Complete and 26 still Ongoing - each a patient cohort united by a common disease, image modality or research focus. Cohort sizes run from dozens of subjects to 26,254 in NLST, and named collections include TCGA-LUAD, CBIS-DDSM, PSMA-PET-CT-Lesions, BraTS-PEDs, HANCOCK and UTSW-Glioma.

What imaging modalities and data types does TCIA cover?

Radiology dominates: CT, MR and PET series appear across the prostate, brain and lung-cohort examples, with mammography present in screening collections. Digital pathology takes the other shape - whole slide images and tissue microarrays, as in HANCOCK. Alongside the pixels, cohorts attach demographic, diagnosis, treatment, molecular-test and segmentation tables where the underlying program collected them.

How far down does the granularity go?

Four levels: patient, imaging study, DICOM series, image instance - with per-collection clinical and genomic tables keyed to the same identifiers. A segmentation model trains at instance level, a cohort study stops at patient level, and both pull from the same schema without reshaping.

Is TCIA suitable for training machine learning models?

Yes - it exists for exactly that. Roughly 240 de-identified cohorts carry expert segmentations, outcome labels and genomics joined to the imagery, so vision models train on pairs rather than on an annotation project. Mind the selection bias: cohorts enter through research programs, so labels reflect what was collected, not population rates.

Can I track changes to collections over time?

Partially, by design. Every Browse Collections row carries an Updated date and a Status flag, and 26 collections are still Ongoing - so growth is observable. But there is no archive-wide version history, so longitudinal views of a specific cohort are built from scheduled captures of the same keyed schema. Because field names never change between captures, snapshots concatenate cleanly.

Which datasets pair well with The Cancer Imaging Archive?

MIMIC-IV supplies tabular intensive-care records for the same patient population question asked without pixels, and PhysioNet adds waveform signal corpora when the study widens beyond imaging. ONC Health IT Dashboard answers the opposite shape - state-year health-IT adoption percentages - which is why we run the head-to-head comparison: pixels versus places, people versus states.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing