Longtian Qiu
发表论文 15 篇 · 总被引 1065 次 · h-index 11
代表论文
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models (2023 · arXiv.org · 被引 309)
- CALIP: Zero-Shot Enhancement of CLIP with Parameter-free Attention (2022 · AAAI Conference on Artificial Intelligence · 被引 196)
- SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models (2024 · International Conference on Machine Learning · 被引 160)
- Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (2024 · arXiv.org · 被引 145)
- HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language Models (2023 · Computer Vision and Pattern Recognition · 被引 99)
- A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise (2023 · arXiv.org · 被引 67)