Y. Shao
发表论文 68 篇 · 总被引 4836 次 · h-index 28
代表论文
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization (2024 · Neural Information Processing Systems · 被引 623)
- RETROSPECTIVE: Aladdin: a Pre-RTL, Power-Performance Accelerator Simulator Enabling Large Design Space Exploration of Customized Architectures (2023 · 被引 185)
- Full Stack Optimization of Transformer Inference: a Survey (2023 · arXiv.org · 被引 176)
- CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale Systems (2023 · International Symposium on Computer Architecture · 被引 40)
- DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators (2023 · Micro · 被引 32)
- AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant Workloads (2023 · Micro · 被引 23)