AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research Paper • 2608.11216 • Published 24 days ago • 2
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published 1 day ago • 2
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization Paper • 2608.12314 • Published 1 day ago • 21
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Paper • 2608.11079 • Published 2 days ago • 12
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 2 days ago • 7
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows Paper • 2608.06714 • Published 6 days ago • 9
Characterizing the Quality Profile of AI-Generated C++ in Production Paper • 2608.06640 • Published 7 days ago • 10
Modular TTT: Rethinking Test-Time Training as Composable Modules Paper • 2608.07110 • Published 6 days ago • 8
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 7 days ago • 46
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 7 days ago • 34
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Paper • 2608.06374 • Published 7 days ago • 23
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published 8 days ago • 58
HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published 8 days ago • 38
OPD-V: Visual On-Policy Self-Distillation with Modality Balance Paper • 2608.05131 • Published 7 days ago • 12
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 9 days ago • 33