FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving Paper • 2502.20238 • Published Feb 27, 2025 • 23
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models Paper • 2403.12027 • Published Mar 18, 2024 • 1
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning Paper • 2509.17437 • Published Sep 22, 2025 • 17
Scaling Environments for LLM Agents in the Era of Learning from Interaction: A Survey Paper • 2511.09586 • Published Nov 12, 2025 • 2
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia Paper • 2511.01670 • Published Nov 3, 2025
Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation Paper • 2406.19643 • Published Jan 3, 2025
Understanding the Behaviors of Environment-aware Information Retrieval Paper • 2606.16817 • Published Jun 15 • 9
Understanding the Behaviors of Environment-aware Information Retrieval Paper • 2606.16817 • Published Jun 15 • 9
In-Context Reinforcement Learning for Tool Use in Large Language Models Paper • 2603.08068 • Published Mar 9 • 43
In-Context Reinforcement Learning for Tool Use in Large Language Models Paper • 2603.08068 • Published Mar 9 • 43
Large Language Models Do NOT Really Know What They Don't Know Paper • 2510.09033 • Published Oct 10, 2025 • 17
Scaling Language-Centric Omnimodal Representation Learning Paper • 2510.11693 • Published Oct 13, 2025 • 108
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use Paper • 2509.24002 • Published Sep 28, 2025 • 180
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Paper • 2507.22607 • Published Jul 30, 2025 • 47
Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations Paper • 2504.13816 • Published Apr 18, 2025 • 18
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning Paper • 2506.07044 • Published Jun 8, 2025 • 114
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning Paper • 2410.22995 • Published Oct 30, 2024 • 3
Babel Collection Open Multilingual Large Language Models Serving Over 90% of Global Speakers • 5 items • Updated Apr 15, 2025 • 18