Norman Mu
发表论文 9 篇 · 总被引 1510 次 · h-index 8
代表论文
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal (2024 · International Conference on Machine Learning · 被引 1448)
- MARKMyWORDS: Analyzing and Evaluating Language Model Watermarks (2023 · 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) · 被引 81)
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs (2025 · Neural Information Processing Systems · 被引 68)
- PAL: Proxy-Guided Black-Box Attack on Large Language Models (2024 · arXiv.org · 被引 58)
- Can LLMs Follow Simple Rules? (2023 · arXiv.org · 被引 54)
- A Closer Look at System Prompt Robustness (2025 · arXiv.org · 被引 38)