toread
updated
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper
• 2604.15574
• Published • 26
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper
• 2604.24763
• Published • 71
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper
• 2604.24819
• Published • 92
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper
• 2604.26752
• Published • 116
Large Language Models Explore by Latent Distilling
Paper
• 2604.24927
• Published • 74
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
Paper
• 2604.26779
• Published • 16
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Paper
• 2605.22791
• Published • 34
Unsupervised Process Reward Models
Paper
• 2605.10158
• Published • 28
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
Paper
• 2605.16928
• Published • 100
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
Paper
• 2605.20177
• Published • 10
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
Paper
• 2605.19282
• Published • 9
Channel-wise Vector Quantization
Paper
• 2605.26089
• Published • 15
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
Paper
• 2605.26895
• Published • 24
Task-Focused Memorization for Multimodal Agents
Paper
• 2605.31075
• Published • 41
LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards
Paper
• 2605.31584
• Published • 43
Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
Paper
• 2605.26844
• Published • 26
ESPO: Early-Stopping Proximal Policy Optimization
Paper
• 2605.29860
• Published • 21
NITP: Next Implicit Token Prediction for LLM Pre-training
Paper
• 2605.24956
• Published • 37
Self-Distilled Policy Gradient
Paper
• 2606.04036
• Published • 28
MemTrain: Self-Supervised Context Memory Training
Paper
• 2606.03197
• Published • 19
Latent Reasoning with Normalizing Flows
Paper
• 2606.06447
• Published • 8
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
Paper
• 2606.03503
• Published • 25
Unified Neural Scaling Laws
Paper
• 2605.26248
• Published • 8
Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
Paper
• 2606.07502
• Published • 100
On the Geometry of On-Policy Distillation
Paper
• 2606.07082
• Published • 75
Paper
• 2606.10650
• Published • 10
How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs
Paper
• 2606.10646
• Published • 9
Rethinking the Divergence Regularization in LLM RL
Paper
• 2606.09821
• Published • 34
Redesign Mixture-of-Experts Routers with Manifold Power Iteration
Paper
• 2606.12397
• Published • 91
Paper
• 2606.13392
• Published • 166
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Paper
• 2606.15007
• Published • 20
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
Paper
• 2606.11176
• Published • 133
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Paper
• 2606.12370
• Published • 22
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
Paper
• 2606.15079
• Published • 88
ExpRL: Exploratory RL for LLM Mid-Training
Paper
• 2606.17024
• Published • 7
Variable-Width Transformers
Paper
• 2606.18246
• Published • 16
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
Paper
• 2606.18216
• Published • 65
Rethinking the Role of Efficient Attention in Hybrid Architectures
Paper
• 2606.15378
• Published • 20
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models
Paper
• 2606.19750
• Published • 4
Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning
Paper
• 2606.18831
• Published • 8
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
Paper
• 2606.24937
• Published • 22
Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning
Paper
• 2606.24133
• Published • 12
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Paper
• 2606.27378
• Published • 61
MultiHashFormer: Hash-based Generative Language Models
Paper
• 2606.28057
• Published • 23
One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining
Paper
• 2606.30634
• Published • 25
Discretizing Reward Models
Paper
• 2606.21795
• Published • 17
Morphing into Hybrid Attention Models
Paper
• 2606.30562
• Published • 52
The State-Prediction Separation Hypothesis
Paper
• 2607.01218
• Published • 12
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Paper
• 2607.02512
• Published • 309
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Paper
• 2607.04439
• Published • 63
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Paper
• 2607.04438
• Published • 64
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Paper
• 2607.04605
• Published • 24
MANCE: Manifold Aware Concept Erasure
Paper
• 2607.03973
• Published • 49
LLM-as-a-Verifier: A General-Purpose Verification Framework
Paper
• 2607.05391
• Published • 18
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper
• 2607.02980
• Published • 84
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Paper
• 2607.04033
• Published • 77
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper
• 2607.05804
• Published • 20
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
Paper
• 2607.07386
• Published • 13
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
Paper
• 2607.07740
• Published • 25
Trust Region Policy Distillation
Paper
• 2607.04751
• Published • 36
Scalable Visual Pretraining for Language Intelligence
Paper
• 2607.09657
• Published • 58
Weak-to-Strong Generalization via Direct On-Policy Distillation
Paper
• 2607.05394
• Published • 150
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
Paper
• 2607.17524
• Published • 7
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Paper
• 2607.16617
• Published • 145
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Paper
• 2607.19691
• Published • 6
Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy Distillation
Paper
• 2607.07050
• Published • 5
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Paper
• 2607.14952
• Published • 213
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Paper
• 2607.21653
• Published • 32
Three-Body Scattering for Generative Modeling
Paper
• 2607.18198
• Published • 16
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
Paper
• 2607.24720
• Published • 26
Kimi K3: Open Frontier Intelligence
Paper
• 2607.24653
• Published • 515
Memory for Large Language Models
Paper
• 2607.25380
• Published • 16
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
Paper
• 2608.20061
• Published • 46
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
Paper
• 2608.31046
• Published • 143
SHAPE of Chain-of-Thought in Math Reasoning
Paper
• 2608.28600
• Published • 31
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Paper
• 2608.30320
• Published • 53
Verification-Aware Training for Speculative Decoding
Paper
• 2608.30135
• Published • 9
Normalized Low-Rank Adaptation
Paper
• 2608.31036
• Published • 50
Language Models Can Control Their Own Attention
Paper
• 2609.02737
• Published • 64
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
Paper
• 2609.02849
• Published • 9
Rethinking On-Policy Distillation of Large Language Models II: One Training Example
Paper
• 2609.04172
• Published • 81
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Paper
• 2608.29188
• Published • 10