Building a Pan-Cancer Atlas of Therapeutic T Cell Targets

How NYU researchers used large-scale immunopeptidomics and AI-enabled peptide identification to identify 28,446 tumor-specific antigens and accelerate therapeutic target discovery.
Institution
NYU Langone Health
Publication
A Pan-Cancer Atlas of Therapeutic T Cell Targets (2025)
Read time
10 min
Lead Researcher
Guangyuan (Frank) Li, PhD, NYU Langone Health

Study at a Glance

2,500+
Raw files processed
<5 Hrs
Analysis completed with Tesorai Search
~60%
More peptide identifications on average
28,446
Tumor-specific HLA-presented antigens identified
ON THIS PAGE

The Challenge

Limited peptide recovery: Traditional database search approaches leave many spectra unidentified, reducing the number of candidate antigens available for downstream analysis.

Operational scalability: Processing thousands of mass spectrometry files through conventional pipelines required extensive hands-on effort and turnaround times measured in weeks or months, making large-scale iterative analysis difficult.

The NYU team needed to:

  • Analyze more than 1,700 immunopeptidomics datasets at scale
  • Search against canonical and non-canonical human proteins ("the dark proteome")
  • Increase peptide identification sensitivity
  • Reduce turnaround times from weeks to hours
  • Enable systematic discovery across 21 cancer types

The Approach

The ImmunoVerse pipeline integrates transcriptomic, immunopeptidomic, and normal tissue datasets to identify tumor-specific antigens arising from multiple classes of molecular alterations.

To maximize peptide recovery from immunopeptidomics data, the team incorporated Tesorai Search into their workflow.

Peptide identification boost: Compared with the baseline search approach used in the study, Tesorai increased peptide identifications by approximately 60% on average. When analysis was restricted to high-confidence peptide-spectrum matches, peptide identifications increased by 136% on average.

At the same time, datasets that previously required weeks of processing through existing workflows could be completed in hours, enabling rapid iteration at scale.

Venn diagram showing peptide overlap between Tesorai Search and other engines in the ImmunoVerse NBL cohort
Figure: Expanded Peptide Discovery with Tesorai Search. Example comparison from the Neuroblastoma (NBL) cohort in the ImmunoVerse study. Tesorai Search identified a large shared set of peptides recovered by other approaches while also identifying additional unique peptides, expanding the pool of candidate antigens available for downstream analysis.

Results

Using Tesorai Search, the NYU team substantially expanded peptide identification across their immunopeptidomics datasets while reducing analysis turnaround times from weeks to hours.

The resulting ImmunoVerse atlas identified:

  • 28,446 tumor-specific HLA-presented antigens
  • 5,928 previously uncharacterized neoantigens
  • 62 previously unrecognized surface protein targets
  • 153 high-priority therapeutic targets

The resource revealed actionable therapeutic targets in 89% of tumors analyzed across 21 cancer types, creating one of the most comprehensive catalogs of cancer-specific antigens reported to date.

"We used to wait days to process thousands of files. Tesorai finished in just a few hours. It completely changes the pace at which we can operate."
Guangyuan (Frank) Li, PhD
NYU Langone Health

Impact

Large-scale immunopeptidomics studies have historically been limited by both analytical sensitivity and computational scalability.

By combining multimodal cancer datasets with AI-enabled peptide identification, the NYU team was able to build ImmunoVerse—a comprehensive atlas of therapeutic T-cell targets spanning 21 cancer types.

The resulting resource provides researchers with a practical foundation for:

  • Cancer vaccine target discovery
  • TCR and T-cell therapy development
  • Tumor antigen prioritization
  • Rapid validation of candidate therapeutic targets

The study demonstrates how improvements in peptide identification and large-scale data processing can transform datasets that were previously too large and complex to analyze efficiently.

Researchers can now use ImmunoVerse together with Tesorai to rapidly evaluate candidate targets against a large repository of immunopeptidomics data without maintaining the underlying computational infrastructure themselves.

About ImmunoVerse

ImmunoVerse is a multimodal cancer antigen atlas integrating transcriptomic, immunopeptidomic, and single-cell datasets across 21 cancer types. The resource is designed to accelerate therapeutic target discovery for cancer vaccines, TCR therapies, and other immune-based treatments.

Published: A Pan-Cancer Atlas of Therapeutic T Cell Targets (2025)

Li G, Yarmarkovich M, et al. | NYU Langone Health

See how Tesorai Search compares

Tesorai identifies up to 68% more peptides than MaxQuant, FragPipe, PEAKS, and Proteome Discoverer — at the same 1% FDR, without match-between-runs.

Explore more resources

Technical explainers, publications, guides, and perspectives on proteomics data analysis and biological discovery.
View all resources

FDR, Match-Between-Runs, and Immunopeptidomics: A Plain-Language Guide

Three concepts that determine whether you can trust your proteomics results — and what to look for when evaluating any search engine.

How Tesorai Works

How Tesorai Search gets more identifications without cutting corners on error control — the scoring model, the FDR methodology, and how it handles both DDA and DIA, in one place.

What is Proteomics? A guide to Mass Spectrometry, Search Engines, and Data Acquisition

From how proteins are measured to why the software that analyzes the data matters as mush as the instrument that generates it.