Jérémy Scheurer
发表论文 16 篇 · 总被引 720 次 · h-index 7
代表论文
- Frontier Models are Capable of In-context Scheming (2024 · arXiv.org · 被引 302)
- Black-Box Access is Insufficient for Rigorous AI Audits (2024 · Conference on Fairness, Accountability and Transparency · 被引 210)
- Large Language Models can Strategically Deceive their Users when Put Under Pressure (2023 · 被引 168)
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs (2024 · arXiv.org · 被引 82)
- Stress Testing Deliberative Alignment for Anti-Scheming Training (2025 · arXiv.org · 被引 66)
- Towards evaluations-based safety cases for AI scheming (2024 · arXiv.org · 被引 39)