WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents Paper • 2609.36887 • Published 12 days ago • 26
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 10 days ago • 282
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Paper • 2610.03574 • Published 9 days ago • 60
DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation Paper • 2610.00360 • Published 11 days ago • 7
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation Paper • 2610.00348 • Published 12 days ago • 14
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs Paper • 2609.32259 • Published 12 days ago • 101
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 12 days ago • 102