LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures Paper • 2610.04292 • Published 9 days ago • 38
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI Paper • 2609.38143 • Published 13 days ago • 85
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published Sep 3 • 188
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published Aug 25 • 152
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 212