Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Protein complexes identification based on go attributed network embedding

作者:Bo Xu, Kun Li, Wei Zheng, Xiaoxia Liu, Yijia Zhang, Zhehuan Zhao, Zengyou He · 发表于:BMC Bioinformatics · 年份:2018 · DOI:10.1186/s12859-018-2555-x · 被引用次数:34 · 研究领域:Bioinformatics and Genomic Networks、Machine Learning in Bioinformatics、Biomedical Text Mining and Ontologies

BACKGROUND: Identifying protein complexes from protein-protein interaction (PPI) network is one of the most important tasks in proteomics. Existing computational methods try to incorporate a variety of biological evidences to enhance the quality of predicted complexes. However, it is still a challenge to integrate different types of biological information into the complexes discovery process under a unified framework. Recently, attributed network embedding methods have be proved to be remarkably effective in generating vector representations for nodes in the network. In the transformed vector space, both the topological proximity and node attributed affinity between different nodes are preserved. Therefore, such attributed network embedding methods provide us a unified framework to integrate various biological evidences into the protein complexes identification process. RESULTS: In this article, we propose a new method called GANE to predict protein complexes based on Gene Ontology (GO) attributed network embedding. Firstly, it learns the vector representation for each protein from a GO attributed PPI network. Based on the pair-wise vector representation similarity, a weighted adjacency matrix is constructed. Secondly, it uses the clique mining method to generate candidate cores. Consequently, seed cores are obtained by ranking candidate cores based on their densities on the weighted adjacency matrix and removing redundant cores. For each seed core, its attachments are the pr...