ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published 8 days ago • 39
Self-Supervised Learning of Structured Dynamics from Videos Paper • 2607.21576 • Published 22 days ago • 20
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 26 days ago • 167
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition Paper • 2601.16211 • Published Jul 2 • 54
Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts Paper • 2607.00666 • Published Jul 1 • 25
Rethinking RAG in Long Videos: What to Retrieve and How to Use It? Paper • 2606.13141 • Published Jun 11 • 36
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding Paper • 2604.05015 • Published Apr 6 • 236
Does Your Reasoning Model Implicitly Know When to Stop Thinking? Paper • 2602.08354 • Published Feb 9 • 267
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models Paper • 2602.04208 • Published Feb 4 • 21
Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance Paper • 2601.14171 • Published Jan 20 • 53