Coleman Hooper
发表论文 34 篇 · 总被引 2214 次 · h-index 15
代表论文
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization (2024 · Neural Information Processing Systems · 被引 673)
- AI and Memory Wall (2024 · IEEE Micro · 被引 413)
- SqueezeLLM: Dense-and-Sparse Quantization (2023 · International Conference on Machine Learning · 被引 365)
- Full Stack Optimization of Transformer Inference: a Survey (2023 · arXiv.org · 被引 178)
- 22.9 A 12nm 18.1TFLOPs/W Sparse Transformer Processor with Entropy-Based Early Exit, Mixed-Precision Predication and Fine-Grained Power Management (2023 · IEEE International Solid-State Circuits Conference · 被引 74)
- TinyAgent: Function Calling at the Edge (2024 · Conference on Empirical Methods in Natural Language Processing · 被引 63)