J. Steinhardt
发表论文 101 篇 · 总被引 37916 次 · h-index 49
代表论文
- Jailbroken: How Does LLM Safety Training Fail? (2023 · Neural Information Processing Systems · 被引 2066)
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small (2022 · International Conference on Learning Representations · 被引 1188)
- Progress measures for grokking via mechanistic interpretability (2023 · International Conference on Learning Representations · 被引 936)
- Discovering Latent Knowledge in Language Models Without Supervision (2022 · International Conference on Learning Representations · 被引 801)
- Scaling Out-of-Distribution Detection for Real-World Settings (2022 · International Conference on Machine Learning · 被引 694)
- Eliciting Latent Predictions from Transformers with the Tuned Lens (2023 · arXiv.org · 被引 549)