Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Y. Shao

发表论文 68 篇 · 总被引 4836 次 · h-index 28

代表论文

  • KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization (2024 · Neural Information Processing Systems · 被引 623)
  • RETROSPECTIVE: Aladdin: a Pre-RTL, Power-Performance Accelerator Simulator Enabling Large Design Space Exploration of Customized Architectures (2023 · 被引 185)
  • Full Stack Optimization of Transformer Inference: a Survey (2023 · arXiv.org · 被引 176)
  • CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale Systems (2023 · International Symposium on Computer Architecture · 被引 40)
  • DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators (2023 · Micro · 被引 32)
  • AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant Workloads (2023 · Micro · 被引 23)