Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Paper • 2512.05774 • Published Dec 5, 2025 • 7
Learning Visual Grounding from Generative Vision and Language Model Paper • 2407.14563 • Published Jul 18, 2024
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Paper • 2607.05390 • Published Jul 6 • 11
A Cookbook of 3D Vision: Data, Learning Paradigms, and Application Paper • 2606.04291 • Published Jun 2 • 6
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use Paper • 2603.08262 • Published Mar 9 • 42
Causal-JEPA: Learning World Models through Object-Level Latent Interventions Paper • 2602.11389 • Published Feb 11 • 13
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models Paper • 2601.05376 • Published Jan 8 • 1
NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos Paper • 2510.08568 • Published Oct 9, 2025 • 2
HyperTaxel: Hyper-Resolution for Taxel-Based Tactile Signals Through Contrastive Learning Paper • 2408.08312 • Published Aug 15, 2024
MotiF: Making Text Count in Image Animation with Motion Focal Loss Paper • 2412.16153 • Published Dec 20, 2024 • 6