Dan Hendrycks
机构:UC Berkeley
发表论文 74 篇 · 总被引 54071 次 · h-index 45
代表论文
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal (2024 · International Conference on Machine Learning · 被引 1401)
- Representation Engineering: A Top-Down Approach to AI Transparency (2023 · arXiv.org · 被引 1228)
- DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models (2023 · Neural Information Processing Systems · 被引 690)
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning (2024 · International Conference on Machine Learning · 被引 501)
- OpenOOD: Benchmarking Generalized Out-of-Distribution Detection (2022 · Neural Information Processing Systems · 被引 391)
- AI deception: A survey of examples, risks, and potential solutions (2023 · Patterns · 被引 377)