Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Long Phan

发表论文 14 篇 · 总被引 412 次 · h-index 8

代表论文

  • Tamper-Resistant Safeguards for Open-Weight LLMs (2024 · International Conference on Learning Representations · 被引 142)
  • Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? (2024 · Neural Information Processing Systems · 被引 82)
  • Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs (2025 · Neural Information Processing Systems · 被引 68)
  • Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State Intervention (2025 · Annual Meeting of the Association for Computational Linguistics · 被引 41)
  • Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark (2025 · arXiv.org · 被引 40)
  • A Definition of AGI (2025 · arXiv.org · 被引 34)