Kyle Montgomery
机构:Washington University in St. Louis · 主页:https://kylemontgomery1.github.io/ · ORCID:0009-0004-0563-347X
发表论文 22 篇 · 总被引 660 次 · h-index 5
代表论文
- A benchmark of expert-level academic questions to assess AI capabilities (2025 · Nature · 被引 392)
- JudgeBench: A Benchmark for Evaluating LLM-based Judges (2024 · International Conference on Learning Representations · 被引 345)
- Agent Instructs Large Language Models to be General Zero-Shot Reasoners (2023 · arXiv.org · 被引 46)
- A Framework for Formalizing LLM Agent Security (2026 · arXiv.org · 被引 17)
- Agents' Last Exam (2026 · arXiv.org · 被引 14)
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess (2025 · arXiv.org · 被引 12)