Linjie Li
机构:Microsoft
发表论文 85 篇 · 总被引 16575 次 · h-index 41
代表论文
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent (2024 · Computer Vision and Pattern Recognition · 被引 241)
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark (2025 · International Conference on Machine Learning · 被引 147)
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement (2025 · Neural Information Processing Systems · 被引 146)
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning (2025 · arXiv.org · 被引 144)
- MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities (2024 · arXiv.org · 被引 64)
- Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback (2025 · International Conference on Machine Learning · 被引 53)