Cross-System, Multi-Scale Single-Cell Foundation Model
with Clinical Applications
A context-aware single-cell transformer that links molecular and cellular variation to patient-level phenotypes across multiple diseases — outperforming all baselines on 50 clinically relevant prediction tasks across 6.4 million cells from 922 donors.
CASCADE learns context-dependent representations across disease, tissue, cell type, and treatment axes, connecting molecular variation at the single-cell level to multiscale biological and clinical phenotypes — from neuropathological staging to treatment response and patient stratification.
Each cell is encoded as context-dependent up- and down-regulated genes relative to a biologically defined reference group — producing multiple representations of the same cell across disease, tissue, cell type, and treatment axes.
Benchmarked across 10 biological and clinical domains spanning neuropathology, genomics, cognition, outcomes, and diagnosis — ranking first in 90.9% of clinical-domain comparisons with a mean improvement of +0.194 over the strongest baseline.
Pre-trained on six disease-specific cohorts spanning Alzheimer's disease, Huntington's disease, lung cancer, lung conditions, autism, and thyroid hormone perturbation — 6.4 million cells from 922 donors.
CASCADE-Explainer identifies cell-type and gene-level drivers of patient predictions. Validated against LLM-based literature rankings (Spearman ρ = 0.964) and AD GWAS-eQTL colocalisation across four independent genetic studies.
CASCADE is pre-trained and evaluated on large-scale single-cell datasets spanning neurodegenerative disease, cancer, and treatment perturbation.
CASCADE integrates contextual information into both input representation and pre-training objectives, allowing the same cell to be interpreted through multiple biologically meaningful axes and enabling patient-level phenotype prediction from single-cell profiles.
Each cell is encoded as context-dependent up- and down-regulated genes relative to a biologically defined reference group, producing multiple representations per cell across disease, tissue, cell type, and treatment contexts.
Shared cell embeddings are projected through separate context-specific projectors (disease, tissue, cell type, treatment), learning how molecular programmes vary across biologically meaningful contexts via contrastive objectives.
Cell-level embeddings are aggregated across all cells from a donor to produce a patient-level representation for multiscale phenotype prediction. CASCADE-Explainer identifies the cell types and genes most responsible for each prediction.
CASCADE was benchmarked against eight baselines — Geneformer, scGPT, scVI, UCE, PaScient, mcBERT, a linear cross-attention model, and a majority-class baseline — across 50 supervised prediction tasks organised into 10 biological and clinical domains. CASCADE ranked first in 90.9% of clinical-domain comparisons, with a mean improvement of +0.194 over the strongest per-task baseline. The largest gains were in Neuropathology & Pathological Staging, Cognitive & Functional Scales, and Outcomes & Cause of Death, where CASCADE ranked first on every task.
CASCADE-Explainer derives donor-specific cell-type and gene-level importance scores from AD status prediction, validated against two independent external benchmarks: a genetic benchmark based on eQTL colocalisation with GWAS loci, and an LLM-arena adjudicating pairwise cell-type comparisons against AD literature. These validations confirm that CASCADE-Explainer captures biologically meaningful disease-relevant signals rather than statistical artefacts.
To assess biological validity, CASCADE-Explainer cell-type importance scores derived from AD status prediction were compared against two independent external benchmarks. The first was an LLM-as-a-judge arena, where pairwise cell-type comparisons were adjudicated using an AD-focused literature report and aggregated into Elo-based rankings — CASCADE-Explainer correlated strongly at coarse resolution (Spearman ρ = 0.964, p = 0.0005) and significantly at granular resolution (ρ = 0.645, p = 0.032). The second was a genetic benchmark based on cell-type-resolved eQTL colocalisation with AD GWAS loci from four independent studies. Positive correlations were observed across all four sources at granular resolution (Bellenguez ρ = 0.387, Jansen ρ = 0.453, Kunkle ρ = 0.567, Marioni ρ = 0.629), strengthening further at coarse resolution (Marioni ρ = 0.775, p = 0.041; Kunkle ρ = 0.764, p = 0.046).
To evaluate gene-level explainability, CASCADE-Explainer was applied to healthy AD donors for a cell-type annotation task and benchmarked against in-house single-nucleus marker programmes (Imperial College London) and literature-derived signatures (Kaufmann et al.). CASCADE-Explainer achieved strong and significant recovery for five non-neuronal cell types — astrocytes, endothelial cells, microglia, OPCs, and oligodendrocytes — with AUROCs ranging from 0.682 to 0.952 (permutation p < 0.001 for all five), using both reference sets. Critically, raw attention weights were near-random (AUROC 0.441–0.570) across all comparisons, confirming that the perturbation-based explainer — not attention alone — drives biologically meaningful gene-level signals.
CASCADE-HD was pre-trained on snRNA-seq data from Huntington's disease cases and healthy controls, then fine-tuned to predict Vonsattel (VS) neuropathological grade, pathogenic CAG repeat length, and benign CAG repeat length — substantially outperforming a regularised linear baseline. CASCADE-Explainer links cell-type programmes to disease severity and enables molecular patient stratification.
CASCADE-Explainer was applied to 53 healthy controls and 41 HD donors across seven major brain cell populations. A clear stage-dependent pattern emerged: spiny projection neuron (SPN) importance increases monotonically from healthy controls through VS Grade 4, reaching its global maximum at the most severe stage. In contrast, microglia show the highest importance in healthy tissue and decline with disease progression. Interneurons peak earliest (VS Grade 1), astrocytes at VS Grade 2, and microglia, oligodendrocytes, and OPCs at VS Grade 3. Endothelial and SPN signals co-peak at VS Grade 4, suggesting that vascular and SPN-associated programmes become most relevant in end-stage neuropathology. These patterns were not driven by cell abundance, confirming that CASCADE-Explainer captures disease-relevant biological variation.
Donor-level cell-type importance profiles from CASCADE-Explainer were used to cluster HD patients in both the benign and pathogenic CAG prediction settings. In both cases, two clusters emerged. In the benign CAG setting, one cluster (N = 22) showed elevated SPN importance and higher VS grade (mean 3.00 vs 1.50), with younger age at death (55.5 vs 66.3 years) and earlier cognitive onset (32.9 vs 57.9 years). The second cluster (N = 28) showed broader non-SPN contributions from microglia, oligodendroglia, and endothelial cells. In the pathogenic CAG setting, the SPN-dominant cluster (N = 37) had longer repeat expansions (mean 46.0 vs 39.8 repeats), higher VS grade, and earlier motor onset (40.4 vs 67.2 years). The benign CAG model captured disease-relevant cellular variability even though benign CAG length itself was not correlated with VS grade, suggesting CASCADE-Explainer identifies non-linear pathology-associated programmes beyond phenotype correlation.
CASCADE-thyroid was pre-trained on mouse cortical scRNA-seq data spanning T3-treated and untreated conditions and wild-type versus dominant-negative thyroid hormone receptor (DN-THRα) perturbations, using cell type and treatment as biological contexts. CASCADE-Explainer was then used to derive gene importance rankings and benchmarked against four curated external reference gene sets spanning transcriptional response and receptor-regulatory biology.
For the treatment task, CASCADE achieved F1 = 0.97 and AUROC = 0.99 (cell-type context) and F1 = 1.00, AUROC = 1.00 (treatment context). For the DN-THRα receptor task, CASCADE achieved F1 = 0.78, AUROC = 0.86 (cell-type context) and F1 = 0.70, AUROC = 0.79 (treatment context). At the cell-type level, CASCADE-Explainer outperformed the differential expression baseline for both astrocytes (AUROC 0.706 vs 0.530, Δ+0.176) and glutamatergic neurons (0.679 vs 0.616, Δ+0.063). The advantage for astrocytes was robust across all aggregation methods tested, reaching a maximum AUROC of 0.745.
A clear double dissociation emerged between the two tasks. For specific TH-responsive genes, CASCADE excelled in the treatment task (+0.095 AUROC over DE in treatment-perturbation), while the receptor task was less advantaged. For genes with proximal THR binding sites, gains were confined to receptor-task settings (+0.050 and +0.037 AUROC), confirming that the receptor task recovered a receptor-proximal regulatory programme rather than a generic transcriptional signature. For broad TH-responsive and DN-THRα de-repressed gene sets, CASCADE consistently outperformed DE across all task–perturbation combinations, with AUROC improvements of +0.071 to +0.228 and +0.047 to +0.122 respectively. Hit-set enrichment further confirmed these results: at Top-500, Fisher odds ratios reached 8.97–12.27 for CASCADE versus 0.99–2.06 for DE, demonstrating that CASCADE-Explainer captures both broad hormone-responsive programmes and receptor-specific regulatory signals.
If you use CASCADE in your research, please cite our paper.
@article{cascade2026, title = {TODO — Paper Title}, author = {TODO — Author, First and Author, Second}, journal = {TODO — Journal or arXiv}, year = {2026}, url = {https://arxiv.org/abs/TODO} }