Xizhou Zhu
发表论文 83 篇 · 总被引 30104 次 · h-index 48
代表论文
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models (2025 · arXiv.org · 被引 1760)
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency (2025 · arXiv.org · 被引 1335)
- Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications (2024 · Computer Vision and Pattern Recognition · 被引 230)
- Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures (2024 · International Conference on Learning Representations · 被引 158)
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning (2025 · arXiv.org · 被引 132)
- The All-Seeing Project V2: Towards General Relation Comprehension of the Open World (2024 · European Conference on Computer Vision · 被引 107)