Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Developing Hybrid Machine Learning Frameworks for Polymer Property Prediction Based on Composition and Sequence Features

作者:Quan Li, Siqi Zhan, Zhanjie Liu, Caibo Dong, Hengheng Zhao, Tongkui Yue, Qingsong Zhao, Liqun Zhang, Ying Li, Jun Liu · 发表于:Journal of Chemical Information and Modeling · 年份:2025 · DOI:10.1021/acs.jcim.5c00745 · 被引用次数:7 · 研究领域:Machine Learning in Materials Science、Computational Drug Discovery Methods

Artificial intelligence (AI) plays a significant role in advancing polymer science and engineering. Considering the critical role of the glass transition temperature ( T g ) in determining the physical properties of polymers, this study systematically investigates the influence of their composition and sequence structure on T g using machine learning (ML) models. To clarify the complex relationship between polymer composition and T g, the k-nearest neighbor mega-trend diffusion (kNNMTD) method was employed for data augmentation, and various ML models were constructed for T g prediction. Among them, the Random Forest model demonstrated the best performance for the generated data, achieving an R 2 of 0.85 and an RMSE of 0.38. To explore the effect of polymer sequence structure on T g, we further introduced natural language processing (NLP) techniques to represent polymer sequences. The data was augmented using the Wasserstein generative adversarial network (GAN) with gradient penalty (WGAN-GP) model, and T g predictions were made using a convolutional neural network-long short-term memory (CNN-LSTM) model. This integrated framework achieved excellent predictive performance, with an R 2 of 0.95 and an RMSE of 0.23, and demonstrated strong generalization across different data sets. In summary, this study introduces an innovative application of kNNMTD for augmenting polymer composition data combined with NLP techniques for representing polymer sequences. The proposed ML framework ...