Cross-modal transfer learning for mapping bulk transcriptomes at cellular level
Abstract
Bulk transcriptomics is widely used in clinical research, yet existing methods struggle to extract single-cell-level structure from bulk measurements, which aggregate signals across heterogeneous cell populations. Here we introduce POPPY, a framework which uses ontology-based contrastive learning to align single-cell and bulk transcriptomic foundation models and construct cell-type-aware bulk patient embeddings. Trained on 1,458 single-cell tumor samples from the Curated Cancer Cell Atlas (3CA) and 1,286 bulk profiles derived from sorted cell populations, POPPY recovers cell type and gene program structure from bulk tumor transcriptomes without fine-tuning. Using a single-cell melanoma atlas, POPPY predicts immunotherapy response across six bulk melanoma cohorts and identifies individual cells associated with response in bulk tumors. These results demonstrate single-cell foundation models can be leveraged to build bulk embeddings that capture cellular biology, enabling interpretable patient stratification from bulk data.