Scholay

学术搜索 · AI 审稿 · LaTeX 协作

KnowCoder-X: Boosting Multilingual Information Extraction via Code

作者:Yuxin Zuo, Wenxuan Jiang, Wenxuan Liu, Zixuan Li, Long Bai, Hanbin Wang, Yutao Zeng, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng · 年份:2025 · DOI:10.18653/v1/2025.findings-acl.748 · 被引用次数:5 · 研究领域:Data Quality and Management、Natural Language Processing Techniques、Advanced Text Analysis Techniques

Empirical evidence indicates that LLMs exhibit spontaneous cross-lingual alignment.However, although LLMs show promising cross-lingual alignment in Information Extraction (IE), a significant imbalance across languages persists, highlighting an underlying deficiency.To address this, we propose KnowCoder-X, a powerful code LLM with advanced cross-lingual and multilingual capabilities for universal IE.Firstly, it standardizes the representation of multilingual schemas using Python classes, ensuring a consistent ontology across different languages.Then, IE across languages is formulated as a unified code generation task.Secondly, we conduct IE cross-lingual alignment instruction tuning on the translated instance prediction task to enhance the model's crosslingual transferability.During this phase, we also construct a high-quality and diverse bilingual IE parallel dataset with 257k samples, called ParallelNER, synthesized by our proposed robust three-stage pipeline, with manual annotation to ensure quality.Although without training in 29 unseen languages, KnowCoder-X surpasses ChatGPT by 30.17% and SoTA by 20.03%, thereby demonstrating superior crosslingual IE capabilities.Comprehensive evaluations on 64 IE benchmarks in Chinese and English under various settings demonstrate that KnowCoder-X significantly enhances crosslingual IE transfer through boosting the IE alignment.Our code and dataset are available at: https://github.com/ICT-GoKnow/ KnowCoder.