Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Prateek Mittal

发表论文 37 篇 · 总被引 2911 次 · h-index 19

代表论文

  • Safety Alignment Should Be Made More Than Just a Few Tokens Deep (2024 · International Conference on Learning Representations · 被引 456)
  • Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications (2024 · International Conference on Machine Learning · 被引 243)
  • SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors (2024 · International Conference on Learning Representations · 被引 227)
  • Certifiably Robust RAG against Retrieval Corruption (2024 · arXiv.org · 被引 142)
  • Data Shapley in One Training Run (2024 · International Conference on Learning Representations · 被引 84)
  • Effectively Controlling Reasoning Models through Thinking Intervention (2025 · arXiv.org · 被引 65)