Dan Braun
发表论文 7 篇 · 总被引 158 次 · h-index 6
代表论文
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning (2024 · Neural Information Processing Systems · 被引 76)
- Towards evaluations-based safety cases for AI scheming (2024 · arXiv.org · 被引 39)
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition (2025 · arXiv.org · 被引 24)
- A Causal Framework for AI Regulation and Auditing (被引 18)
- Stochastic Parameter Decomposition (2025 · arXiv.org · 被引 14)
- Using Degeneracy in the Loss Landscape for Mechanistic Interpretability (2024 · arXiv.org · 被引 14)