Building a Pan-Cancer Atlas of Therapeutic T Cell Targets
Study at a Glance
The Challenge
Limited peptide recovery: Traditional database search approaches leave many spectra unidentified, reducing the number of candidate antigens available for downstream analysis.
Operational scalability: Processing thousands of mass spectrometry files through conventional pipelines required extensive hands-on effort and turnaround times measured in weeks or months, making large-scale iterative analysis difficult.
The NYU team needed to:
- Analyze more than 1,700 immunopeptidomics datasets at scale
- Search against canonical and non-canonical human proteins ("the dark proteome")
- Increase peptide identification sensitivity
- Reduce turnaround times from weeks to hours
- Enable systematic discovery across 21 cancer types
The Approach
The ImmunoVerse pipeline integrates transcriptomic, immunopeptidomic, and normal tissue datasets to identify tumor-specific antigens arising from multiple classes of molecular alterations.
To maximize peptide recovery from immunopeptidomics data, the team incorporated Tesorai Search into their workflow.
Peptide identification boost: Compared with the baseline search approach used in the study, Tesorai increased peptide identifications by approximately 60% on average. When analysis was restricted to high-confidence peptide-spectrum matches, peptide identifications increased by 136% on average.
At the same time, datasets that previously required weeks of processing through existing workflows could be completed in hours, enabling rapid iteration at scale.

Results
Using Tesorai Search, the NYU team substantially expanded peptide identification across their immunopeptidomics datasets while reducing analysis turnaround times from weeks to hours.
The resulting ImmunoVerse atlas identified:
- 28,446 tumor-specific HLA-presented antigens
- 5,928 previously uncharacterized neoantigens
- 62 previously unrecognized surface protein targets
- 153 high-priority therapeutic targets
The resource revealed actionable therapeutic targets in 89% of tumors analyzed across 21 cancer types, creating one of the most comprehensive catalogs of cancer-specific antigens reported to date.
"We used to wait days to process thousands of files. Tesorai finished in just a few hours. It completely changes the pace at which we can operate."
Guangyuan (Frank) Li, PhD
NYU Langone Health
Impact
Large-scale immunopeptidomics studies have historically been limited by both analytical sensitivity and computational scalability.
By combining multimodal cancer datasets with AI-enabled peptide identification, the NYU team was able to build ImmunoVerse—a comprehensive atlas of therapeutic T-cell targets spanning 21 cancer types.
The resulting resource provides researchers with a practical foundation for:
- Cancer vaccine target discovery
- TCR and T-cell therapy development
- Tumor antigen prioritization
- Rapid validation of candidate therapeutic targets
The study demonstrates how improvements in peptide identification and large-scale data processing can transform datasets that were previously too large and complex to analyze efficiently.
Researchers can now use ImmunoVerse together with Tesorai to rapidly evaluate candidate targets against a large repository of immunopeptidomics data without maintaining the underlying computational infrastructure themselves.
About ImmunoVerse
ImmunoVerse is a multimodal cancer antigen atlas integrating transcriptomic, immunopeptidomic, and single-cell datasets across 21 cancer types. The resource is designed to accelerate therapeutic target discovery for cancer vaccines, TCR therapies, and other immune-based treatments.
Published: A Pan-Cancer Atlas of Therapeutic T Cell Targets (2025)
Li G, Yarmarkovich M, et al. | NYU Langone Health
