Xuwang Yin
发表论文 8 篇 · 总被引 2453 次 · h-index 5
代表论文
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal (2024 · International Conference on Machine Learning · 被引 1448)
- Representation Engineering: A Top-Down Approach to AI Transparency (2023 · arXiv.org · 被引 1275)
- Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? (2024 · Neural Information Processing Systems · 被引 82)
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs (2025 · Neural Information Processing Systems · 被引 68)
- The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems (2025 · arXiv.org · 被引 50)
- Jailbreak Distillation: Renewable Safety Benchmarking (2025 · Conference on Empirical Methods in Natural Language Processing · 被引 4)