Saurav Kadavath
机构:UC Berkeley
发表论文 41 篇 · 总被引 21634 次 · h-index 16
代表论文
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback (2022 · arXiv.org · 被引 4293)
- Constitutional AI: Harmlessness from AI Feedback (2022 · arXiv.org · 被引 3506)
- Language Models (Mostly) Know What They Know (2022 · arXiv.org · 被引 1855)
- Measuring Coding Challenge Competence With APPS (2021 · NeurIPS Datasets and Benchmarks · 被引 1234)
- Discovering Language Model Behaviors with Model-Written Evaluations (2022 · Annual Meeting of the Association for Computational Linguistics · 被引 967)
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned (2022 · arXiv.org · 被引 859)