Publication

Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling

Fine-tuning a single-cell foundation model on proteomics data matches or beats scaling to models over 40x larger — showing multimodal training can outperform parameter scaling alone.
ON THIS PAGE

Tesorai's research team has posted a new preprint on bioRxiv showing that adding proteomics data to single-cell foundation model training can outperform simply scaling up model size. By fine-tuning a 70M-parameter Tahoe-x1 model on proteomic profiles from 440 diverse mass-spectrometry studies, the team matched or exceeded the performance of 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 benchmarks.

Read the paper: https://doi.org/10.64898/2026.08.14.744845 

Abstract

Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.

Explore more resources

Technical explainers, publications, guides, and perspectives on proteomics data analysis and biological discovery.
View all resources
Publication

Publications

Tesorai Search: Cloud-based Database Search Engine Boosts Identifications for Mass Spectrometry Proteomics With a Pretrained Peptide-spectrum Deep-learning Model

Tesorai Search is now peer-reviewed and published in the Journal of Molecular Biology, detailing the model's design and benchmarking against other leading proteomics search engines.

Scientific poster

Publications

Rapid Mechanistic Bridging of an Alzheimer's Disease Plasma Protein Staging Panel Across Brain Proteomics Cohorts

Five public brain proteomics cohorts, reprocessed on Tesorai, support a seven-protein blood panel for Alzheimer's staging, including recovered p-tau217. Presented at HUPO 2026.

Scientific poster

Publications

Developing and Benchmarking Agentic LLM Frameworks for Accelerated and Accessible Proteomics Data Analysis

On a 94-case proteomics benchmark, Tesorai Chat reached 92.2% accuracy, ahead of Phylo, GPT-5.4 Codex, and Claude 4.6 Sonnet. Presented at HUPO 2026.