Zhenkai Liang
发表论文 13 篇 · 总被引 85 次 · h-index 5
代表论文
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint (2025 · arXiv.org · 被引 40)
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards (2025 · arXiv.org · 被引 17)
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning (2025 · arXiv.org · 被引 12)
- Your Scale Factors are My Weapon: Targeted Bit-Flip Attacks on Vision Transformers via Scale Factor Manipulation (2025 · Computer Vision and Pattern Recognition · 被引 10)
- DARWIN: An approach to debugging evolving programs (2012 · TSEM · 被引 6)
- PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants (2025 · arXiv.org · 被引 4)