Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Bo An

发表论文 11 篇 · 总被引 433 次 · h-index 6

代表论文

  • Group-in-Group Policy Optimization for LLM Agent Training (2025 · Neural Information Processing Systems · 被引 413)
  • SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning (2025 · arXiv.org · 被引 147)
  • Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks (2026 · arXiv.org · 被引 33)
  • Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds (2025 · arXiv.org · 被引 16)
  • Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning (2025 · International Conference on Machine Learning · 被引 15)
  • Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems (2026 · arXiv.org · 被引 12)