Structured Residual Connectivity Matters for Diffusion Transformers Paper • 2609.33203 • Published 3 days ago • 15
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published Jul 16 • 45
dOPSD: On-Policy Self-Distillation for Diffusion Language Models Paper • 2607.04428 • Published Jul 5 • 18
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Paper • 2605.20315 • Published May 19 • 26
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment Paper • 2605.11882 • Published May 12 • 15
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Paper • 2604.28185 • Published Apr 30 • 90
Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms Paper • 2604.23775 • Published Apr 26 • 45
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers Paper • 2603.27666 • Published Mar 29 • 16
Can MLLMs Guide Me Home? A Benchmark Study on Fine-Grained Visual Reasoning from Transit Maps Paper • 2505.18675 • Published May 24, 2025 • 29
Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models Paper • 2603.15557 • Published Mar 16 • 30
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer Paper • 2603.15478 • Published Mar 16 • 24