Scholay

学术搜索 · AI 审稿 · LaTeX 协作

AgentMol: Multi-Model AI System for Automatic Drug-Target Identification and Molecule Development

作者:Piotr Karabowicz, R. Charkiewicz, A. Charkiewicz, Anetta Sulewska, Jacek Nikliński · 发表于:Methods and Protocols · 年份:2025 · DOI:10.3390/mps8060143 · 被引用次数:2 · 研究领域:Medicine

Drug discovery remains a time-consuming and costly process, necessitating innovative computational approaches to accelerate early stage target identification and compound development. We introduce AgentMol, a modular multimodel AI system that integrates large language models, chemical language modeling, and deep learning–based affinity prediction to automate the discovery pipeline. AgentMol begins with disease-related queries processed through a Retrieval-Augmented Generation system using the Large Language Model to identify protein targets. Protein sequences are then used to condition a GPT-2–based chemical language model, which generates corresponding small-molecule candidates in SMILES format. Finally, a regression convolutional neural network (RCNN) predicts the drug-target interaction by estimating binding affinities (pKi). Models were trained and validated on 470,560 ligand–protein pairs from the BindingDB database. The chemical language model achieved high validity (1.00), uniqueness (0.96), and diversity (0.89), whereas the RCNN model demonstrated robust predictive performance with R2 > 0.6 and Pearson’s R > 0.8. By leveraging LangGraph for orchestration, AgentMol delivers a scalable, interpretable pipeline, effectively enabling the end-to-end generation and evaluation of drug candidates conditioned on protein targets. This system represents a significant step toward practical AI-driven molecular discovery with accessible computational demands.