VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction Paper • 2610.06293 • Published 6 days ago • 55
Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 176
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published Sep 5 • 166
SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models Paper • 2608.29098 • Published Aug 29 • 10
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 52
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published Jul 21 • 114