FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 2 days ago • 94
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 2 days ago • 94
AutoForge: Automated Environment Synthesis for Agentic Reinforcement Learning Paper • 2512.22857 • Published Dec 28, 2025 • 1
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome Paper • 2603.28407 • Published Mar 30 • 72
AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents Paper • 2603.27490 • Published Mar 29 • 20
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 3 days ago • 188
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 3 days ago • 188
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published 16 days ago • 64