WorldGuide: Goal-Directed Video World Model for Procedural Task Execution Paper • 2610.12459 • Published 4 days ago • 8
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation Paper • 2610.01092 • Published 11 days ago • 34
Training-Free Speech-Centric Omni Understanding with Frozen VLMs Paper • 2609.04242 • Published Aug 7 • 7
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published Jul 30 • 62
Next-Embedding Prediction Makes Strong Vision Learners Paper • 2512.16922 • Published Dec 18, 2025 • 91