Abstract
Pretrained models, originally developed for vision and textual data, are not a panacea and may fail to fully represent the complexity of sequences in immunological tasks. In studying pretrained immunological sequence modules of a renowned immunogenicity prediction model, pMTnet, we observe that our carefully designed model removing (ablating) the pretrained T cell receptor (TCR) autoencoder in pMTnet can even improve the prediction accuracy. Furthermore, we note the TCR pretraining data, used by pretrained modules within pMTnet, dramatically deviates from a broader and more representative TCR repertoire. Such findings underscore the impacts of the blend of heterogeneous representations and distribution discrepancy in immunological sequences, which necessitate appropriate coordination of different pretrained models and representative databases.
| Original language | English |
|---|---|
| Number of pages | 17 |
| Journal | Communications Biology |
| DOIs | |
| Publication status | E-pub ahead of print - 20 Jul 2026 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Fingerprint
Dive into the research topics of 'Pretrained models may fail to capture immunological sequences'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver