Mikita Balesni
发表论文 19 篇 · 总被引 1379 次 · h-index 11
代表论文
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" (2023 · International Conference on Learning Representations · 被引 535)
- Frontier Models are Capable of In-context Scheming (2024 · arXiv.org · 被引 302)
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (2025 · arXiv.org · 被引 225)
- Large Language Models can Strategically Deceive their Users when Put Under Pressure (2023 · 被引 168)
- Taken out of context: On measuring situational awareness in LLMs (2023 · arXiv.org · 被引 147)
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs (2024 · arXiv.org · 被引 82)