Nicholas Joseph
发表论文 26 篇 · 总被引 25080 次 · h-index 18
代表论文
- Evaluating Large Language Models Trained on Code (2021 · arXiv.org · 被引 11018)
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback (2022 · arXiv.org · 被引 4303)
- Constitutional AI: Harmlessness from AI Feedback (2022 · arXiv.org · 被引 3538)
- Language Models (Mostly) Know What They Know (2022 · arXiv.org · 被引 1877)
- A General Language Assistant as a Laboratory for Alignment (2021 · arXiv.org · 被引 1196)
- In-context Learning and Induction Heads (2022 · arXiv.org · 被引 981)