Hongsheng Li
发表论文 18 篇 · 总被引 1304 次 · h-index 13
代表论文
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models (2023 · arXiv.org · 被引 309)
- SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models (2024 · International Conference on Machine Learning · 被引 160)
- Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining (2024 · arXiv.org · 被引 156)
- Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (2024 · arXiv.org · 被引 145)
- Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want (2024 · International Conference on Learning Representations · 被引 111)
- Lumina-Image 2.0: a Unified and Efficient Image Generative Framework (2025 · IEEE International Conference on Computer Vision · 被引 96)