Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Yongming Rao

机构:Nanyang Technological University, Tencent (China) · ORCID:0000-0003-3952-8753

发表论文 125 篇 · 总被引 7576 次 · h-index 37

代表论文

  • Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models (2025 · 被引 13)
  • Vision Generalist Model: A Survey (2025 · International Journal of Computer Vision · 被引 3)
  • Efficient High-Order Spatial Interactions for Visual Perception (2025 · IEEE Transactions on Pattern Analysis and Machine Intelligence · 被引 2)
  • Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training (2026 · Lecture notes in computer science · 被引 1)
  • Generative Multimodal Models Are In-Context Learners (2025 · Advances in computer vision and pattern recognition · 被引 1)
  • GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots (2026 · arXiv (Cornell University))