Classification of Collagens via Peptide Ambiguation, in a Paleoproteomic LC-MS/MS-Based Taxonomic Pipeline
作者:Ian Engels, Alexandra Burnett, Prudence Robert, Camille Pironneau, Grégory Abrams, Robbin Bouwmeester, P. Van Der Plaetsen, Kévin Di Modica, Marcel Otte, Lawrence Guy Straus, Valentin Fischer, Fabrice Bray, Bart Mesuere, Isabelle De Groote, Dieter L. D. Deforce, Simon Daled, Maarten Dhaenens · 发表于:Journal of Proteome Research · 年份:2025 · DOI:10.1021/acs.jproteome.4c00962 · 被引用次数:13 · 研究领域:Advanced Proteomics Techniques and Applications、Identification and Quantification in Food、Metabolomics and Mass Spectrometry Studies
Liquid chromatography-mass spectrometry (LC-MS/MS) extends the matrix-assisted laser desorption ionization-time of flight (MALDI-TOF) Zooarcheology by Mass Spectrometry (ZooMS) "mass fingerprinting" approach to species identification by providing fragmentation spectra for each peptide. However, ancient bone samples generate sparse data containing only a few collagen proteins, rendering target-decoy strategies unusable and increasing uncertainty in peptide annotation. To ameliorate this issue, we present a ZooMS/MS data pipeline that builds on a manually curated Collagen database and comprises two novel algorithms: isoBLAST and ClassiCOL. isoBLAST first extends peptide ambiguity by generating all "potential peptide candidates" isobaric to the annotated precursor. The exhaustive set of candidates created is then used to retain or reject different potential paths at each taxonomic branching point from superkingdom to species, until the greatest possible specificity is reached. Uniquely, ClassiCOL allows for the identification of taxonomic mixtures, including contaminated samples, as well as suggesting taxonomies not represented in sequence databases, including extinct taxa. All considered ambiguity is then graphically represented with clear prioritization of the potential taxa in the sample. Using public as well as in-house data acquired on different instruments, we demonstrate the performance of this universal postprocessing and explore the identification of both genetic and sa...