Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Junyoung Park

发表论文 12 篇 · 总被引 122 次 · h-index 5

代表论文

  • KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments (2025 · Advances in Neural Information Processing Systems 38 · 被引 37)
  • On Speculative Decoding for Multimodal Large Language Models (2024 · arXiv.org · 被引 34)
  • Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement (2024 · arXiv.org · 被引 29)
  • Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs (2024 · arXiv.org · 被引 19)
  • VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs (2025 · arXiv.org · 被引 10)
  • CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction (2025 · 被引 7)