Scholay

学术搜索 · AI 审稿 · LaTeX 协作

RARE: Retrieval-Augmented Reasoning Modeling

作者:Zhengren Wang, Jiayang Yu, Dongsheng Ma, Zhe Chen, Yu Wang, Zhiyu Li, Feiyu Xiong, Yanfeng Wang, E Weinan, Linpeng Tang, Wentao Zhang · 年份:2026 · DOI:10.1145/3770855.3817899 · 被引用次数:2 · 研究领域:Multimodal Machine Learning Applications、Intelligent Tutoring Systems and Adaptive Learning、Topic Modeling

Domain-specific intelligence demands specialized knowledge and sophisticated reasoning for problem-solving, posing significant challenges for large language models (LLMs) that struggle with knowledge hallucination and inadequate reasoning capabilities. For the efficient discovery and modeling of reasoning patterns, we propose Retrieval-Augmented Reasoning Modeling (RARE), a novel paradigm to extract and model reasoning patterns from data, represented as salient and learnable tokens. Specifically, RARE externalizes domain knowledge to retrievable sources and internalizes reasoning patterns in LLMs. By injecting retrieved knowledge into training prompts with masked losses, RARE transforms learning objectives from rote memorization to contextualized reasoning, making implicit reasoning patterns salient. It enables models to bypass parameter-intensive memorization and prioritize the modeling of higher-order reasoning patterns. Extensive experiments demonstrate that lightweight RARE-trained models (e.g., Llama-3.1-8B and Qwen3-8B) could achieve state-of-the-art performance, even rivaling retrieval-augmented Gemini-3-Pro-Preview and GPT-5.1 with trillion parameters. RARE catalyzes a paradigm shift where maintainable external knowledge bases synergize with compact, reasoning-optimized models, collectively driving more scalable domain-specific intelligence.