Martin T. Vechev
发表论文 311 篇 · 总被引 17053 次 · h-index 67
代表论文
- MathArena: Evaluating LLMs on Uncontaminated Math Competitions (2025 · Advances in Neural Information Processing Systems 38 · 被引 303)
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs (2026 · arXiv.org · 被引 57)
- ToolFuzz - Automated Agent Tool Testing (2025 · arXiv.org · 被引 18)
- Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? (2026 · arXiv.org · 被引 12)
- Coding Agents Don't Know When to Act (2026 · arXiv.org · 被引 4)
- Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness (2026 · arXiv.org · 被引 3)