YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection Paper • 2512.23273 • Published Dec 29, 2025 • 15
A 58-Addition, Rank-23 Scheme for General 3x3 Matrix Multiplication Paper • 2512.21980 • Published Dec 26, 2025 • 3
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding Paper • 2512.16229 • Published Dec 18, 2025 • 17
CASA: Cross-Attention via Self-Attention for Efficient Vision-Language Fusion Paper • 2512.19535 • Published Dec 22, 2025 • 13
Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Paper • 2512.17351 • Published Dec 19, 2025 • 29
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing Paper • 2512.14681 • Published Dec 16, 2025 • 44
Janus: Disaggregating Attention and Experts for Scalable MoE Inference Paper • 2512.13525 • Published Dec 15, 2025 • 6
QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management Paper • 2512.12967 • Published Dec 15, 2025 • 113
Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics Paper • 2512.12602 • Published Dec 14, 2025 • 44
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Paper • 2604.04921 • Published Apr 6 • 117
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference Paper • 2604.07394 • Published Apr 8 • 16
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Paper • 2605.27365 • Published May 26 • 147
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Paper • 2607.07740 • Published Jul 8 • 26
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing Paper • 2607.07953 • Published Jul 8 • 16
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 173
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching Paper • 2608.08097 • Published Aug 8 • 26
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving Paper • 2608.19758 • Published 28 days ago • 20
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 4 days ago • 34