Development of a large pretrained model for Tesorai Search

Peter's lecture at MaxQuant Summer School 2024 describing the principles underlying how Search was built.
ON THIS PAGE

Peter Cimermancic was invited to give a lecture at MaxQuant Summer School, a long-running bootcamp for mass spec proteomics analysis. He talks about outstanding challenges in the field and how Tesorai Search uses AI to overcome them.

Watch his seminar: https://www.youtube.com/watch?v=RMVboOlBaLk 

AI-generated summary of Peter’s seminar

Peter Cimermancic, co-founder of Tesorai, introduces Tesorai Search, a new platform for computational mass spectrometry data analysis. He highlights significant data loss (up to 75% of spectra) during initial pre-processing in mass spectrometry due to older software and overfitting issues in "Gen 2" AI/ML algorithms (0:33-2:25, 6:24-12:02).

Tesorai's solution is a single, end-to-end, entirely pre-trained AI model designed to accurately identify peptides without common overfitting problems (12:03-16:32). This "Gen 3" approach boosts peptide identifications by 5% to 90% across various datasets, shows strong generalizability across different instruments and sample types, and leads to the discovery of significantly more proteins, validated by human experts and increased overlap with known biological interactions (16:33-23:48).

The Tesorai Search platform is a user-friendly, scalable, cloud-based web application that processes hundreds of files in less than an hour (23:53-27:52). Future plans include "Instant Search" for near-real-time results and "Incognito Search" to utilize 100% of spectral data for differential analysis, dramatically increasing biomarker discovery potential (28:19-31:20).

Explore more resources

Technical explainers, publications, guides, and perspectives on proteomics data analysis and biological discovery.
View all resources
Publication

Publications

Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling

Fine-tuning a single-cell foundation model on proteomics data matches or beats scaling to models over 40x larger — showing multimodal training can outperform parameter scaling alone.

Pan-Cancer Atlas resource card graphic

Building a Pan-Cancer Atlas of Therapeutic T Cell Targets

How NYU researchers used large-scale immunopeptidomics and AI-enabled peptide identification to identify 28,446 tumor-specific antigens and accelerate therapeutic target discovery.

Guangyuan (Frank) Li, PhD, NYU Langone Health
FDR, MBR, and Immunopeptidomics guide resource card graphic

FDR, Match-Between-Runs, and Immunopeptidomics: A Plain-Language Guide

Three concepts that determine whether you can trust your proteomics results — and what to look for when evaluating any search engine.