β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 10 days ago • 23
FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse Paper • 2606.11290 • Published Jun 9 • 2
TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models Paper • 2601.18744 • Published Jan 26 • 11
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation Paper • 2511.01163 • Published Nov 3, 2025 • 32
view article Article Jupyter Agents: training LLMs to reason with notebooks +1 baptistecolle, hannayukhymenko, lvwerra • Sep 10, 2025 • 67
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases Paper • 2407.12784 • Published Jul 17, 2024 • 51