Collections
Discover the best community collections!
Collections including paper arxiv:2607.08964
-
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
Paper • 2604.18543 • Published • 30 -
Near-Future Policy Optimization
Paper • 2604.20733 • Published • 77 -
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
Paper • 2604.20987 • Published • 22 -
PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering
Paper • 2602.23161 • Published
-
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
Paper • 2510.03222 • Published • 76 -
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
Paper • 2510.05592 • Published • 113 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 519 -
Multi-Agent Tool-Integrated Policy Optimization
Paper • 2510.04678 • Published • 31
-
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
Paper • 2505.13227 • Published • 46 -
facebook/natural_reasoning
Viewer • Updated • 1.15M • 3.11k • 580 -
nvidia/OpenMathReasoning
Viewer • Updated • 5.68M • 48.7k • 470 -
Search Arena: Analyzing Search-Augmented LLMs
Paper • 2506.05334 • Published • 19
-
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 264 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 106 -
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Paper • 2607.08964 • Published • 77 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 145
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 76 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 19
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 6 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 59 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 51
-
4D Human-Scene Reconstruction from Low-Overlap Captures
Paper • 2607.09125 • Published • 53 -
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
Paper • 2505.24858 • Published • 17 -
Metacognition in LLMs: Foundations, Progress, and Opportunities
Paper • 2607.11881 • Published • 30 -
Video Generation Models are General-Purpose Vision Learners
Paper • 2607.09024 • Published • 87
-
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 264 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 106 -
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Paper • 2607.08964 • Published • 77 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 145
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 76 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 19
-
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
Paper • 2604.18543 • Published • 30 -
Near-Future Policy Optimization
Paper • 2604.20733 • Published • 77 -
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
Paper • 2604.20987 • Published • 22 -
PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering
Paper • 2602.23161 • Published
-
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Paper • 2603.19220 • Published • 70 -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Paper • 2605.20164 • Published • 6 -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Paper • 2605.19577 • Published • 59 -
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
Paper • 2605.18703 • Published • 51
-
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
Paper • 2510.03222 • Published • 76 -
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
Paper • 2510.05592 • Published • 113 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 519 -
Multi-Agent Tool-Integrated Policy Optimization
Paper • 2510.04678 • Published • 31
-
4D Human-Scene Reconstruction from Low-Overlap Captures
Paper • 2607.09125 • Published • 53 -
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
Paper • 2505.24858 • Published • 17 -
Metacognition in LLMs: Foundations, Progress, and Opportunities
Paper • 2607.11881 • Published • 30 -
Video Generation Models are General-Purpose Vision Learners
Paper • 2607.09024 • Published • 87
-
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
Paper • 2505.13227 • Published • 46 -
facebook/natural_reasoning
Viewer • Updated • 1.15M • 3.11k • 580 -
nvidia/OpenMathReasoning
Viewer • Updated • 5.68M • 48.7k • 470 -
Search Arena: Analyzing Search-Augmented LLMs
Paper • 2506.05334 • Published • 19