Song Han
发表论文 21 篇 · 总被引 2283 次 · h-index 13
代表论文
- VILA: On Pre-training for Visual Language Models (2023 · Computer Vision and Pattern Recognition · 被引 908)
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models (2025 · Computer Vision and Pattern Recognition · 被引 508)
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos (2024 · International Conference on Learning Representations · 被引 314)
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation (2024 · International Conference on Learning Representations · 被引 283)
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos (2025 · arXiv.org · 被引 122)
- WorldModelBench: Judging Video Generation Models As World Models (2025 · arXiv.org · 被引 86)