Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation Paper • 2610.05608 • Published 3 days ago • 122
The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 14 days ago • 64
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control Paper • 2609.17909 • Published 22 days ago • 47
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published Jul 18 • 140
EgoCS-400K: An Egocentric Gameplay Dataset for World Models Paper • 2606.18180 • Published Jun 16 • 17
HumanNet: Scaling Human-centric Video Learning to One Million Hours Paper • 2605.06747 • Published May 7 • 55
Pretraining Frame Preservation in Autoregressive Video Memory Compression Paper • 2512.23851 • Published Dec 29, 2025 • 25
LongVie 2: Multimodal Controllable Ultra-Long Video World Model Paper • 2512.13604 • Published Dec 15, 2025 • 76
FlashWorld: High-quality 3D Scene Generation within Seconds Paper • 2510.13678 • Published Oct 15, 2025 • 74
Diffusion Transformers with Representation Autoencoders Paper • 2510.11690 • Published Oct 13, 2025 • 171
SPATIALGEN: Layout-guided 3D Indoor Scene Generation Paper • 2509.14981 • Published Sep 18, 2025 • 28
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding Paper • 2509.15178 • Published Sep 18, 2025 • 6
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels Paper • 2507.21809 • Published Jul 29, 2025 • 142