Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Xuwang Yin

发表论文 8 篇 · 总被引 2453 次 · h-index 5

代表论文

  • HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal (2024 · International Conference on Machine Learning · 被引 1448)
  • Representation Engineering: A Top-Down Approach to AI Transparency (2023 · arXiv.org · 被引 1275)
  • Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? (2024 · Neural Information Processing Systems · 被引 82)
  • Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs (2025 · Neural Information Processing Systems · 被引 68)
  • The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems (2025 · arXiv.org · 被引 50)
  • Jailbreak Distillation: Renewable Safety Benchmarking (2025 · Conference on Empirical Methods in Natural Language Processing · 被引 4)