Buck Shlegeris
发表论文 29 篇 · 总被引 2830 次 · h-index 18
代表论文
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small (2022 · International Conference on Learning Representations · 被引 1275)
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training (2024 · arXiv.org · 被引 565)
- Alignment faking in large language models (2024 · arXiv.org · 被引 337)
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (2025 · arXiv.org · 被引 225)
- AI Control: Improving Safety Despite Intentional Subversion (2023 · International Conference on Machine Learning · 被引 213)
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models (2024 · arXiv.org · 被引 160)