Research Paper  ·  2026

CASCADE

Cross-System, Multi-Scale Single-Cell Foundation Model
with Clinical Applications

A context-aware single-cell transformer that links molecular and cellular variation to patient-level phenotypes across multiple diseases — outperforming all baselines on 50 clinically relevant prediction tasks across 6.4 million cells from 922 donors.

Valentina Giunchiglia Owen Queen Xiang Lin Zaneta Matuszek Daniel Hochbaum Aarthi Venkat Gianmarco Abbadessa Walker Rickord Richard Nicholas Paola Arlotta Marinka Zitnik Michelle Li
In collaboration with
Harvard University Medical Research Council Broad Institute Imperial College London King's College London Stanford University
Overview

Context-aware single-cell modelling
across diseases and scales

CASCADE learns context-dependent representations across disease, tissue, cell type, and treatment axes, connecting molecular variation at the single-cell level to multiscale biological and clinical phenotypes — from neuropathological staging to treatment response and patient stratification.

CASCADE — context-aware single-cell modelling overview
🧬
Context-Aware Tokenisation

Each cell is encoded as context-dependent up- and down-regulated genes relative to a biologically defined reference group — producing multiple representations of the same cell across disease, tissue, cell type, and treatment axes.

50 Clinical Prediction Tasks

Benchmarked across 10 biological and clinical domains spanning neuropathology, genomics, cognition, outcomes, and diagnosis — ranking first in 90.9% of clinical-domain comparisons with a mean improvement of +0.194 over the strongest baseline.

📊
Multi-Disease Scope

Pre-trained on six disease-specific cohorts spanning Alzheimer's disease, Huntington's disease, lung cancer, lung conditions, autism, and thyroid hormone perturbation — 6.4 million cells from 922 donors.

🔬
Biological Explainability

CASCADE-Explainer identifies cell-type and gene-level drivers of patient predictions. Validated against LLM-based literature rankings (Spearman ρ = 0.964) and AD GWAS-eQTL colocalisation across four independent genetic studies.

Dataset

Six disease-specific cohorts

CASCADE is pre-trained and evaluated on large-scale single-cell datasets spanning neurodegenerative disease, cancer, and treatment perturbation.

6.4M+
Total cells
922
Donors
50
Prediction tasks
6
Disease cohorts
Cohorts: Alzheimer's Disease · Huntington's Disease · Lung Cancer · Lung Conditions · Autism · Thyroid Perturbation
Architecture

Three modules bridging cells
to patient phenotypes

CASCADE integrates contextual information into both input representation and pre-training objectives, allowing the same cell to be interpreted through multiple biologically meaningful axes and enabling patient-level phenotype prediction from single-cell profiles.

1
Context-aware tokenisation

Each cell is encoded as context-dependent up- and down-regulated genes relative to a biologically defined reference group, producing multiple representations per cell across disease, tissue, cell type, and treatment contexts.

2
Context-specific representation learning

Shared cell embeddings are projected through separate context-specific projectors (disease, tissue, cell type, treatment), learning how molecular programmes vary across biologically meaningful contexts via contrastive objectives.

3
Patient representation & explainability

Cell-level embeddings are aggregated across all cells from a donor to produce a patient-level representation for multiscale phenotype prediction. CASCADE-Explainer identifies the cell types and genes most responsible for each prediction.

CASCADE architecture
Figure 1c. CASCADE architecture: context-aware tokenisation module, context-specific projectors (cell type, tissue, disease, treatment), patient representation module, and explainability module.
Benchmarking

State-of-the-art across 50
clinical prediction tasks

CASCADE was benchmarked against eight baselines — Geneformer, scGPT, scVI, UCE, PaScient, mcBERT, a linear cross-attention model, and a majority-class baseline — across 50 supervised prediction tasks organised into 10 biological and clinical domains. CASCADE ranked first in 90.9% of clinical-domain comparisons, with a mean improvement of +0.194 over the strongest per-task baseline. The largest gains were in Neuropathology & Pathological Staging, Cognitive & Functional Scales, and Outcomes & Cause of Death, where CASCADE ranked first on every task.

Ranked 1st in 90.9% of clinical-domain comparisons Mean improvement +0.194 over strongest per-task baseline 50 tasks across 4 disease cohorts  ·  922 donors
CASCADE benchmarking across 50 prediction tasks
Figure 2. Benchmarking of CASCADE across 50 clinically relevant prediction tasks. Top: Multi-context shared embedding space — UMAPs coloured by disease, cell-type, and tissue context. Bottom: F1 scores per task domain across four cohorts (HLCA, LUCA, SEATTLE, AUTISM), compared against eight baselines.
Alzheimer's Disease

CASCADE-Explainer identifies cell-type
drivers of Alzheimer's disease

CASCADE-Explainer derives donor-specific cell-type and gene-level importance scores from AD status prediction, validated against two independent external benchmarks: a genetic benchmark based on eQTL colocalisation with GWAS loci, and an LLM-arena adjudicating pairwise cell-type comparisons against AD literature. These validations confirm that CASCADE-Explainer captures biologically meaningful disease-relevant signals rather than statistical artefacts.

Alzheimer's Disease — Explainer Validation

Cell-type importance rankings align with GWAS and literature evidence

To assess biological validity, CASCADE-Explainer cell-type importance scores derived from AD status prediction were compared against two independent external benchmarks. The first was an LLM-as-a-judge arena, where pairwise cell-type comparisons were adjudicated using an AD-focused literature report and aggregated into Elo-based rankings — CASCADE-Explainer correlated strongly at coarse resolution (Spearman ρ = 0.964, p = 0.0005) and significantly at granular resolution (ρ = 0.645, p = 0.032). The second was a genetic benchmark based on cell-type-resolved eQTL colocalisation with AD GWAS loci from four independent studies. Positive correlations were observed across all four sources at granular resolution (Bellenguez ρ = 0.387, Jansen ρ = 0.453, Kunkle ρ = 0.567, Marioni ρ = 0.629), strengthening further at coarse resolution (Marioni ρ = 0.775, p = 0.041; Kunkle ρ = 0.764, p = 0.046).

LLM-arena: ρ = 0.964  (p = 0.0005, coarse cell types) eQTL concordance significant across 4 independent AD GWAS studies
CASCADE-Explainer AD validation — GWAS and LLM benchmarking
Figure 3a–e. Orthogonal benchmarking of CASCADE-Explainer cell-type importance scores. a) Evaluation design: eQTL colocalisation (genetic) and LLM-arena (literature-based). b–c) LLM-arena concordance at granular and coarse cell-type resolution. d–e) eQTL colocalisation concordance at granular and coarse resolution.
Alzheimer's Disease — Gene Programmes

CASCADE-Explainer recovers canonical cell-type marker programmes

To evaluate gene-level explainability, CASCADE-Explainer was applied to healthy AD donors for a cell-type annotation task and benchmarked against in-house single-nucleus marker programmes (Imperial College London) and literature-derived signatures (Kaufmann et al.). CASCADE-Explainer achieved strong and significant recovery for five non-neuronal cell types — astrocytes, endothelial cells, microglia, OPCs, and oligodendrocytes — with AUROCs ranging from 0.682 to 0.952 (permutation p < 0.001 for all five), using both reference sets. Critically, raw attention weights were near-random (AUROC 0.441–0.570) across all comparisons, confirming that the perturbation-based explainer — not attention alone — drives biologically meaningful gene-level signals.

AUROC 0.682–0.952 across 5 non-neuronal cell types  (p < 0.001) Replicated across in-house and literature reference sets Attention baseline near-random (AUROC 0.441–0.570)
CASCADE-Explainer gene programme recovery
Figure 3f. Gene programme recovery AUROCs per cell type (Astrocytes, Endothelial, IO neurons, Microglia, OPCs, Oligodendrocytes). Bars show CASCADE-Explainer, attention baseline, and DE baseline against in-house and literature reference marker sets.
Huntington's Disease

CASCADE predicts neuropathological
severity and enables patient stratification

CASCADE-HD was pre-trained on snRNA-seq data from Huntington's disease cases and healthy controls, then fine-tuned to predict Vonsattel (VS) neuropathological grade, pathogenic CAG repeat length, and benign CAG repeat length — substantially outperforming a regularised linear baseline. CASCADE-Explainer links cell-type programmes to disease severity and enables molecular patient stratification.

CASCADE-HD overview pipeline
CASCADE-HD: patient representations from Huntington's disease and healthy single-cell data are used to predict benign CAG, pathogenic CAG, and VS grade, with donor-level cell-type importance driving patient clustering into clinically distinct subgroups.
Cell-Type Importance by VS Grade

SPN importance tracks neuropathological severity monotonically

CASCADE-Explainer was applied to 53 healthy controls and 41 HD donors across seven major brain cell populations. A clear stage-dependent pattern emerged: spiny projection neuron (SPN) importance increases monotonically from healthy controls through VS Grade 4, reaching its global maximum at the most severe stage. In contrast, microglia show the highest importance in healthy tissue and decline with disease progression. Interneurons peak earliest (VS Grade 1), astrocytes at VS Grade 2, and microglia, oligodendrocytes, and OPCs at VS Grade 3. Endothelial and SPN signals co-peak at VS Grade 4, suggesting that vascular and SPN-associated programmes become most relevant in end-stage neuropathology. These patterns were not driven by cell abundance, confirming that CASCADE-Explainer captures disease-relevant biological variation.

VS Grade: r = 0.863, R² = 0.611 Pathogenic CAG: r = 0.599  ·  Benign CAG: r = 0.502 SPN importance increases monotonically from healthy → VS Grade 4 Microglia show strongest negative association with VS grade (ρ = −0.824)
Cell-type importance by VS grade — SPN and non-neuronal patterns
Left: Observed correlations between predicted HD phenotypes (VS grade, benign and pathogenic CAG) and their Δ biological and model validity. Right: Cohort-level cell-type importance across neuropathological severity stages, showing SPN monotonic increase and non-neuronal cell-type peaks.
Patient Stratification

Cell-type importance profiles stratify patients into clinically distinct clusters

Donor-level cell-type importance profiles from CASCADE-Explainer were used to cluster HD patients in both the benign and pathogenic CAG prediction settings. In both cases, two clusters emerged. In the benign CAG setting, one cluster (N = 22) showed elevated SPN importance and higher VS grade (mean 3.00 vs 1.50), with younger age at death (55.5 vs 66.3 years) and earlier cognitive onset (32.9 vs 57.9 years). The second cluster (N = 28) showed broader non-SPN contributions from microglia, oligodendroglia, and endothelial cells. In the pathogenic CAG setting, the SPN-dominant cluster (N = 37) had longer repeat expansions (mean 46.0 vs 39.8 repeats), higher VS grade, and earlier motor onset (40.4 vs 67.2 years). The benign CAG model captured disease-relevant cellular variability even though benign CAG length itself was not correlated with VS grade, suggesting CASCADE-Explainer identifies non-linear pathology-associated programmes beyond phenotype correlation.

Benign CAG: 2 clusters (silhouette = 0.325), VS grade 3.00 vs 1.50 Pathogenic CAG: 2 clusters (silhouette = 0.588), motor onset 40 vs 67 yrs Clusters validated against age at death, motor and cognitive onset
HD patient clustering by cell-type importance profiles
Panels d–g. Donor-level cell-type importance for benign and pathogenic CAG predictions: forest plots of clinical correlations (d), heatmaps of donor variability per cluster (e), mean VS grade per cluster (f), and pairwise cell-type weight matrices within and between clusters (g).
Thyroid Perturbation

CASCADE-Explainer recovers
thyroid hormone gene signatures

CASCADE-thyroid was pre-trained on mouse cortical scRNA-seq data spanning T3-treated and untreated conditions and wild-type versus dominant-negative thyroid hormone receptor (DN-THRα) perturbations, using cell type and treatment as biological contexts. CASCADE-Explainer was then used to derive gene importance rankings and benchmarked against four curated external reference gene sets spanning transcriptional response and receptor-regulatory biology.

CASCADE-thyroid benchmark overview
Benchmark design: CASCADE-Explainer gene rankings compared against matched DE rankings across two tasks (thyroid hormone treatment; DN-THRα receptor prediction) and four reference gene sets (specific TH-responsive, broad TH-responsive, proximal THR binding sites, DN-THRα de-repressed).
Prediction Performance

CASCADE recovers cell-type-specific thyroid hormone programmes

For the treatment task, CASCADE achieved F1 = 0.97 and AUROC = 0.99 (cell-type context) and F1 = 1.00, AUROC = 1.00 (treatment context). For the DN-THRα receptor task, CASCADE achieved F1 = 0.78, AUROC = 0.86 (cell-type context) and F1 = 0.70, AUROC = 0.79 (treatment context). At the cell-type level, CASCADE-Explainer outperformed the differential expression baseline for both astrocytes (AUROC 0.706 vs 0.530, Δ+0.176) and glutamatergic neurons (0.679 vs 0.616, Δ+0.063). The advantage for astrocytes was robust across all aggregation methods tested, reaching a maximum AUROC of 0.745.

Treatment task: AUROC = 0.99–1.00 Receptor task: AUROC = 0.79–0.86 Astrocytes: CASCADE 0.706 vs DE 0.530 (Δ+0.176) Glutamatergic neurons: CASCADE 0.679 vs DE 0.616 (Δ+0.063)
Aggregation method benchmark and cell-type AUROC
b. Aggregation method benchmark — median AUROC across all task–perturbation–reference combinations. c. Cell-type-specific AUROC for CASCADE vs DE in astrocytes and glutamatergic neurons.
Task-Specific Enrichment

CASCADE captures distinct transcriptional and receptor-regulatory programmes

A clear double dissociation emerged between the two tasks. For specific TH-responsive genes, CASCADE excelled in the treatment task (+0.095 AUROC over DE in treatment-perturbation), while the receptor task was less advantaged. For genes with proximal THR binding sites, gains were confined to receptor-task settings (+0.050 and +0.037 AUROC), confirming that the receptor task recovered a receptor-proximal regulatory programme rather than a generic transcriptional signature. For broad TH-responsive and DN-THRα de-repressed gene sets, CASCADE consistently outperformed DE across all task–perturbation combinations, with AUROC improvements of +0.071 to +0.228 and +0.047 to +0.122 respectively. Hit-set enrichment further confirmed these results: at Top-500, Fisher odds ratios reached 8.97–12.27 for CASCADE versus 0.99–2.06 for DE, demonstrating that CASCADE-Explainer captures both broad hormone-responsive programmes and receptor-specific regulatory signals.

Broad TH-responsive: Δ AUROC +0.071 to +0.228 over DE DN-THRα de-repressed: Δ AUROC +0.047 to +0.122 over DE Top-500 enrichment: odds ratio 8.97–12.27 vs 0.99–2.06 for DE
Thyroid enrichment and AUROC heatmaps
d–e. Fisher exact odds ratios for broad TH-responsive genes and proximal THR binding sites across tasks and perturbation settings. f–g. Full AUROC matrix (CASCADE) and Δ AUROC (CASCADE − DE) across all task–perturbation–reference combinations.
Authors

Research Team

Owen Queen
Owen Queen
Xiang Lin
Xiang Lin
Zaneta Matuszek
Zaneta Matuszek
Daniel Hochbaum
Daniel Hochbaum
Michelle Li
Michelle Li
Aarthi Venkat
Aarthi Venkat
Gianmarco Abbadessa
Gianmarco Abbadessa
Imperial College London
Walker Rickord
Walker Rickord
Richard Nicholas
Richard Nicholas
Paola Arlotta
Paola Arlotta
Marinka Zitnik
Marinka Zitnik
Citation

Cite CASCADE

If you use CASCADE in your research, please cite our paper.

BibTeX
@article{cascade2026,
  title   = {TODO — Paper Title},
  author  = {TODO — Author, First and Author, Second},
  journal = {TODO — Journal or arXiv},
  year    = {2026},
  url     = {https://arxiv.org/abs/TODO}
}