Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Overcoming Shortcut Learning in RNA-Small Molecule Modeling via Bias-Matched Decoys and Structure-Aware Network Design

作者:Yiming Wen, Yilin Han, Dingyan Wang · 发表于:Journal of Chemical Information and Modeling · 年份:2026 · DOI:10.1021/acs.jcim.6c00360 · 被引用次数:2 · 研究领域:Computational Drug Discovery Methods、RNA and protein synthesis mechanisms、RNA Research and Splicing

Targeting RNA with small molecules represents a promising frontier in drug discovery, yet current computational models are often hindered by the physicochemical property misalignment within available training sets. In this study, we reveal that widely used machine learning training configurations, such as the ROBIN training set, exhibit a significant physicochemical mismatch. Specifically, the curated negative samples, which are predominantly sourced from external databases like BindingDB, exhibit prominent property disparities compared to RNA-targeted ligands. These macroscopic differences provide a “shortcut” for models to achieve high predictive accuracy through simple property filtering, thereby overshadowing the genuine structural recognition patterns required for RNA binding. To address this, we introduce RNAdecoyDB, a property-matched hard decoy data set generated through a feature-matching strategy that aligns the physicochemical property distributions between positive samples and decoys. Building upon this unbiased data set, we develop the Target RNA Network (TRN), a ligand-centric predictive model designed to capture intrinsic ligand structural features governing RNA-binding propensity. Unlike models that require target RNA information, TRN operates solely on 2D molecular graphs to identify structural motifs enriched in RNA-binding ligands. Comprehensive cross-distribution benchmarks and external validation on independent data sets (R-SIM and SM2miR) demonstrate tha...