Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Song Han

发表论文 21 篇 · 总被引 2283 次 · h-index 13

代表论文

  • VILA: On Pre-training for Visual Language Models (2023 · Computer Vision and Pattern Recognition · 被引 908)
  • CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models (2025 · Computer Vision and Pattern Recognition · 被引 508)
  • LongVILA: Scaling Long-Context Visual Language Models for Long Videos (2024 · International Conference on Learning Representations · 被引 314)
  • VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation (2024 · International Conference on Learning Representations · 被引 283)
  • EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos (2025 · arXiv.org · 被引 122)
  • WorldModelBench: Judging Video Generation Models As World Models (2025 · arXiv.org · 被引 86)