HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 15 days ago • 339
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 6 days ago • 177
The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents Paper • 2608.06065 • Published 26 days ago • 8
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 22 days ago • 342
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow Paper • 2607.28362 • Published Jul 30 • 22
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 140
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition Paper • 2607.25294 • Published Jul 28 • 49
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published Jul 22 • 193