Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Qwen3-VL Technical Report

作者:Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Deng, Lianghao, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Jiang, Shutong, Zhaohai Li, Mingsheng Li, Minyong Li, Keqiang Li, Zicheng Lin, Lin, Junyang, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yan Liu, Dayiheng Liu, Shixuan Liu, Dan Lü, Ruilin Luo, Chenxu Lv, Rui Men, Lingchen Meng, Xuancheng Ren, Xingzhang Ren, Sibo Song, Yuchong Sun, Jun Tang, Jianhong Tu, Jianqiang Wan, Peng Wang, Pengfei Wang, Qiuyue Wang, Yuxuan Wang, Tianbao Xie, Yiheng Xu, Haiyang Xu, Jin Xu, Zhibo Yang, Mingkun Yang, Jianxin Yang, Yang An, Bowen Yu, Zhang, Fei, Zhang Hang, Xi Zhang, Bo Zheng, Zhong, Humen, Jingren Zhou, Fan Zhou, Jing Zhou, Yuanzhi Zhu, K. J. Zhu · 发表于:arXiv (Cornell University) · 年份:2025 · DOI:10.48550/arxiv.2511.21631 · 被引用次数:13 · 研究领域:Multimodal Machine Learning Applications、Topic Modeling、Generative Adversarial Networks and Image Synthesis

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL...