ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Paper • 2608.05102 • Published 1 day ago • 47
MiniWorld: Democratizing the Training of Video World Models from Scratch Paper • 2608.01127 • Published 5 days ago • 13
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation Paper • 2608.00079 • Published 9 days ago • 14
ShotPlan: Cinematic Video Generation with Learnable Planning Token Paper • 2607.17675 • Published 18 days ago • 6
DiFA: Inference-Time Forward-Process Alignment for Diffusion Models Paper • 2607.17972 • Published 18 days ago • 7
Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator Paper • 2607.05765 • Published about 1 month ago • 3
CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation Paper • 2607.03803 • Published Jul 4 • 19
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Paper • 2607.06559 • Published about 1 month ago • 95
Unified Audio Intelligence Without Regressing on Text Intelligence Paper • 2607.05196 • Published Jul 6 • 23
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model Paper • 2607.03509 • Published Jul 3 • 14
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
Unified Panoramic Geometry Estimation via Multi-View Foundation Models Paper • 2605.26368 • Published May 25 • 4
OminiControl: Minimal and Universal Control for Diffusion Transformer Paper • 2411.15098 • Published Nov 22, 2024 • 61
POSTA: A Go-to Framework for Customized Artistic Poster Generation Paper • 2503.14908 • Published Mar 19, 2025
OminiControl2: Efficient Conditioning for Diffusion Transformers Paper • 2503.08280 • Published Mar 11, 2025
Master: Meta Style Transformer for Controllable Zero-Shot and Few-Shot Artistic Style Transfer Paper • 2304.11818 • Published Apr 24, 2023