Unveiling potential threats: backdoor attacks in single-cell pre-trained models
作者:Sicheng Feng, Siyu Li, Luonan Chen, Shengquan Chen · 发表于:Cell Discovery · 年份:2024 · DOI:10.1038/s41421-024-00753-1 · 被引用次数:5 · 研究领域:Single-cell and spatial transcriptomics、Adversarial Robustness in Machine Learning、SARS-CoV-2 and COVID-19 Research
The advancement of single-cell sequencing technology has empowered fields such as developmental biology, immunology, and oncology, underscoring its significance in revealing individual cell characteristics in health and disease. Many computational methods and workflows have been specifically designed for single-cell data analysis to accurately characterize cellular heterogeneity 1 . The accumulation of extensive single-cell datasets and the ongoing refinement of comprehensive cell atlases have catalyzed the development of advanced pre-trained models such as scBERT 2 , GeneFormer 3 , and scGPT 4 . These models facilitate versatile downstream analyses, including cell type annotation, gene regulatory network inference, and drug response prediction, outperforming specialized methods tailored for the corresponding tasks 5 . To pre-train such powerful models, the collection of vast amounts of training data is essential. Besides, due to computational resource constraints, training is often outsourced to third parties, or pre-trained models from external sources are utilized. However, stemming from unintentional issues in sample preparation, data processing, or cell type annotation, as well as intentional poisoning driven by commercial interests, single-cell pre-trained models face potential threats of backdoor attacks (Fig. 1a ), which differ from accidental noise (Supplementary Text S 1 ) and can severely impact biomedical research by compromising their integrity and reliability (S...