Yongming Rao
机构:Nanyang Technological University, Tencent (China) · ORCID:0000-0003-3952-8753
发表论文 125 篇 · 总被引 7576 次 · h-index 37
代表论文
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models (2025 · 被引 13)
- Vision Generalist Model: A Survey (2025 · International Journal of Computer Vision · 被引 3)
- Efficient High-Order Spatial Interactions for Visual Perception (2025 · IEEE Transactions on Pattern Analysis and Machine Intelligence · 被引 2)
- Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (2026 · Lecture notes in computer science · 被引 1)
- Generative Multimodal Models Are In-Context Learners (2025 · Advances in computer vision and pattern recognition · 被引 1)
- GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots (2026 · arXiv (Cornell University))