TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 5 days ago • 88
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability Paper • 2610.08448 • Published 5 days ago • 194
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework Paper • 2609.30541 • Published 17 days ago • 3
Agent Plasticity: Measuring Self-Improvement Through Experience Paper • 2610.08902 • Published 5 days ago • 5
Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute Paper • 2610.10845 • Published 4 days ago • 8
UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy Paper • 2610.10164 • Published 4 days ago • 7
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles Paper • 2610.09832 • Published 4 days ago • 8
Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position Paper • 2610.10114 • Published 4 days ago • 31
ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation Paper • 2609.39306 • Published 11 days ago • 33
SGF+: Decoupling Gradient Flows for Autoregressive Video Generation Paper • 2610.10429 • Published 4 days ago • 55
Questioning the Questions: Sustaining Self-Evolution in Reasoning Models Paper • 2610.04299 • Published 8 days ago • 70
GRACE: Generation-aware latent compression for efficient video generation Paper • 2610.10524 • Published 4 days ago • 76
Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness Paper • 2610.08621 • Published 5 days ago • 87
nanoMuse: An Open-Source Personal Agent for Every Device You Own Paper • 2610.08699 • Published 5 days ago • 101
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization Paper • 2609.38169 • Published 12 days ago • 117