Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Finding the dark matter: Large language model‐based enzyme kinetic data extractor and its validation

作者:Galen Wei, Xinchun Ran, Runeem AI‐Abssi, Zhongyue Yang · 发表于:Protein Science · 年份:2025 · DOI:10.1002/pro.70251 · 被引用次数:6 · 研究领域:Bioinformatics and Genomic Networks、Machine Learning in Bioinformatics、Protein Structure and Dynamics

Abstract Despite the vast number of enzymatic kinetic measurements reported across decades of biochemical literature, the majority of relational enzyme kinetic data—linking amino acid sequence, substrate identity, kinetic parameters, and assay conditions—remains uncollected and inaccessible in structured form. This constitutes a significant portion of the “dark matter” of enzymology. Unlocking these hidden data through automated extraction offers an opportunity to expand enzyme dataset diversity and size, critical for building accurate, generalizable models that drive predictive enzyme engineering. To address this limitation, we built EnzyExtract, a large language model‐powered pipeline that automates the extraction, verification, and structuring of enzyme kinetics data from scientific literature. By processing 137,892 full‐text publications (PDF/XML), EnzyExtract collected more than 218,095 enzyme–substrate–kinetics entries, including 218,095 k cat and 167,794 K m values. These entries are mapped to enzymes spanning 3569 unique four‐digit EC numbers, with a total of 84,464 entries assigned at least a first‐digit EC number. EnzyExtract identified 89,544 unique kinetic entries ( k cat and K m combined) absent from BRENDA, significantly expanding the known enzymology dataset. The newly curated dataset was compiled into a database named EnzyExtractDB. EnzyExtract demonstrates high accuracy when benchmarked against manually curated datasets and strong consistency with BRENDA‐deri...