Long Phan
发表论文 14 篇 · 总被引 412 次 · h-index 8
代表论文
- Tamper-Resistant Safeguards for Open-Weight LLMs (2024 · International Conference on Learning Representations · 被引 142)
- Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? (2024 · Neural Information Processing Systems · 被引 82)
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs (2025 · Neural Information Processing Systems · 被引 68)
- Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State Intervention (2025 · Annual Meeting of the Association for Computational Linguistics · 被引 41)
- Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark (2025 · arXiv.org · 被引 40)
- A Definition of AGI (2025 · arXiv.org · 被引 34)