Saining Xie
发表论文 64 篇 · 总被引 7387 次 · h-index 29
代表论文
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset (2025 · arXiv.org · 被引 367)
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers (2025 · IEEE International Conference on Computer Vision · 被引 204)
- Meta CLIP 2: A Worldwide Scaling Recipe (2025 · Neural Information Processing Systems · 被引 67)
- What matters for Representation Alignment: Global Information or Spatial Structure? (2025 · arXiv.org · 被引 67)
- Scaling Inference Time Compute for Diffusion Models (2025 · Computer Vision and Pattern Recognition · 被引 51)
- Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders (2026 · arXiv.org · 被引 44)