Alexander Meinke
发表论文 8 篇 · 总被引 389 次 · h-index 5
代表论文
- Frontier Models are Capable of In-context Scheming (2024 · arXiv.org · 被引 302)
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs (2024 · arXiv.org · 被引 82)
- Stress Testing Deliberative Alignment for Anti-Scheming Training (2025 · arXiv.org · 被引 66)
- Towards evaluations-based safety cases for AI scheming (2024 · arXiv.org · 被引 39)
- Towards a Situational Awareness Benchmark for LLMs (被引 11)
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs (2024 · Advances in Neural Information Processing Systems 37 · 被引 3)