Marius Hobbhahn
发表论文 20 篇 · 总被引 1008 次 · h-index 12
代表论文
- Frontier Models are Capable of In-context Scheming (2024 · arXiv.org · 被引 303)
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (2025 · arXiv.org · 被引 224)
- Black-Box Access is Insufficient for Rigorous AI Audits (2024 · Conference on Fairness, Accountability and Transparency · 被引 210)
- Large Language Models can Strategically Deceive their Users when Put Under Pressure (2023 · 被引 168)
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs (2024 · arXiv.org · 被引 82)
- Detecting Strategic Deception Using Linear Probes (2025 · arXiv.org · 被引 69)