Prateek Mittal
发表论文 37 篇 · 总被引 2911 次 · h-index 19
代表论文
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep (2024 · International Conference on Learning Representations · 被引 456)
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications (2024 · International Conference on Machine Learning · 被引 243)
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors (2024 · International Conference on Learning Representations · 被引 227)
- Certifiably Robust RAG against Retrieval Corruption (2024 · arXiv.org · 被引 142)
- Data Shapley in One Training Run (2024 · International Conference on Learning Representations · 被引 84)
- Effectively Controlling Reasoning Models through Thinking Intervention (2025 · arXiv.org · 被引 65)