Publication

Tesorai Search: Cloud-based Database Search Engine Boosts Identifications for Mass Spectrometry Proteomics With a Pretrained Peptide-spectrum Deep-learning Model

Tesorai Search is now peer-reviewed and published in the Journal of Molecular Biology, detailing the model's design and benchmarking against other leading proteomics search engines.
ON THIS PAGE

Tesorai Search is now peer-reviewed and published in the Journal of Molecular Biology, describing the implementation of Tesorai Search, including detailed benchmarking of Tesorai against other leading software across multiple datasets.

Read the paper: https://doi.org/10.1016/j.jmb.2026.169940 

Abstract

The original mass spectrometry search engines used simple algorithms for peptide identification. Recent tools improved accuracy by adding several extra components such as fragment ion intensities or retention times prediction and training target-decoy classifiers on-the-fly, leading to sometimes inconsistent results.

Our study explores the impact of replacing those extra components with a deep-learning pretrained model that directly learns the complex relationship between the full spectra and associated peptide sequence, without using decoys. This simplified workflow has fewer parameters to tweak, making it easier to use and perform robustly on data from instruments and use-cases never seen during training.

Surprisingly, our approach consistently identifies more peptides than FragPipe, PEAKS, and Proteome Discoverer (12%, 9%, and 21% more, respectively, across a range of datasets). Tesorai Search is also fast – 250 immunopeptidomics searches in 45 minutes – and free for academics, available as a webserver at console.tesorai.com.

Explore more resources

Technical explainers, publications, guides, and perspectives on proteomics data analysis and biological discovery.
View all resources
Publication

Publications

Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling

Fine-tuning a single-cell foundation model on proteomics data matches or beats scaling to models over 40x larger — showing multimodal training can outperform parameter scaling alone.

Pan-Cancer Atlas resource card graphic

Building a Pan-Cancer Atlas of Therapeutic T Cell Targets

How NYU researchers used large-scale immunopeptidomics and AI-enabled peptide identification to identify 28,446 tumor-specific antigens and accelerate therapeutic target discovery.

Guangyuan (Frank) Li, PhD, NYU Langone Health
FDR, MBR, and Immunopeptidomics guide resource card graphic

FDR, Match-Between-Runs, and Immunopeptidomics: A Plain-Language Guide

Three concepts that determine whether you can trust your proteomics results — and what to look for when evaluating any search engine.