Nicholas Goldowsky-Dill
发表论文 10 篇 · 总被引 515 次 · h-index 7
代表论文
- Localizing Model Behavior with Path Patching (2023 · arXiv.org · 被引 202)
- Open Problems in Mechanistic Interpretability (2025 · arXiv.org · 被引 198)
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning (2024 · Neural Information Processing Systems · 被引 76)
- Detecting Strategic Deception Using Linear Probes (2025 · arXiv.org · 被引 69)
- Towards evaluations-based safety cases for AI scheming (2024 · arXiv.org · 被引 39)
- Detecting Strategic Deception with Linear Probes (2025 · International Conference on Machine Learning · 被引 30)