JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published 25 days ago • 69
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 106
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Paper • 2607.28509 • Published Jul 30 • 30
Beacon: Knowing When and How to Perform Agentic Visual Reasoning Paper • 2607.28595 • Published Jul 30 • 56
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Paper • 2607.14202 • Published Jul 15 • 44
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Paper • 2607.14189 • Published Jul 15 • 31
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Paper • 2607.07608 • Published Jul 8 • 57
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Paper • 2607.07608 • Published Jul 8 • 57
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation Paper • 2606.17628 • Published Jun 16 • 30
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 218
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Paper • 2605.05997 • Published May 7 • 18