Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Lijie, Zhang, Zhihao, Jain, Arti, Cao, Shijie, Yuan, Baihong, Chen, Yiwei, Jia, Zhihao, Netravali, Ravi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
LIMO: Less is More for Reasoning
von: Ye, Yixin, et al.
Veröffentlicht: (2025)
von: Ye, Yixin, et al.
Veröffentlicht: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024)
von: Lou, Chao, et al.
Veröffentlicht: (2024)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
von: Zhuang, Xialie, et al.
Veröffentlicht: (2025)
von: Zhuang, Xialie, et al.
Veröffentlicht: (2025)
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
von: Zhang, Chenyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyuan, et al.
Veröffentlicht: (2026)
Accelerating Retrieval-Augmented Language Model Serving with Speculation
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
Rectified Sparse Attention
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
Geometry Guided Self-Consistency for Physical AI
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
A Unified Sparse Attention via Multi-Granularity Compression
von: Liu, Siran, et al.
Veröffentlicht: (2025)
von: Liu, Siran, et al.
Veröffentlicht: (2025)
Acting Less is Reasoning More! Teaching Model to Act Efficiently
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
von: Hoang, Duy C., et al.
Veröffentlicht: (2024)
von: Hoang, Duy C., et al.
Veröffentlicht: (2024)
Less Data Less Tokens: Multilingual Unification Learning for Efficient Test-Time Reasoning in LLMs
von: Chen, Kang, et al.
Veröffentlicht: (2025)
von: Chen, Kang, et al.
Veröffentlicht: (2025)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
von: Yu, Jiachen, et al.
Veröffentlicht: (2026)
von: Yu, Jiachen, et al.
Veröffentlicht: (2026)
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
von: Zhu, Jinchang, et al.
Veröffentlicht: (2026)
von: Zhu, Jinchang, et al.
Veröffentlicht: (2026)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
von: Liu, Jinzhe, et al.
Veröffentlicht: (2025)
von: Liu, Jinzhe, et al.
Veröffentlicht: (2025)
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation
von: Ray, Siddhant, et al.
Veröffentlicht: (2024)
von: Ray, Siddhant, et al.
Veröffentlicht: (2024)
Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning
von: Sui, Yi, et al.
Veröffentlicht: (2026)
von: Sui, Yi, et al.
Veröffentlicht: (2026)
Head-wise Shareable Attention for Large Language Models
von: Cao, Zouying, et al.
Veröffentlicht: (2024)
von: Cao, Zouying, et al.
Veröffentlicht: (2024)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
von: Liu, Siran, et al.
Veröffentlicht: (2026)
von: Liu, Siran, et al.
Veröffentlicht: (2026)
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
von: Zhang, Qianchi, et al.
Veröffentlicht: (2025)
von: Zhang, Qianchi, et al.
Veröffentlicht: (2025)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Soft Prompt Tuning for Cross-Lingual Transfer: When Less is More
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
LLMSR@XLLM25: Less is More: Enhancing Structured Multi-Agent Reasoning via Quality-Guided Distillation
von: Yuan, Jiahao, et al.
Veröffentlicht: (2025)
von: Yuan, Jiahao, et al.
Veröffentlicht: (2025)
SIBO: A Simple Booster for Parameter-Efficient Fine-Tuning
von: Wen, Zhihao, et al.
Veröffentlicht: (2024)
von: Wen, Zhihao, et al.
Veröffentlicht: (2024)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
von: Hu, Junhao, et al.
Veröffentlicht: (2026)
von: Hu, Junhao, et al.
Veröffentlicht: (2026)
BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2025)
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2025)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
von: Chen, Si, et al.
Veröffentlicht: (2024)
von: Chen, Si, et al.
Veröffentlicht: (2024)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2024) -
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
von: Pan, Rui, et al.
Veröffentlicht: (2025) -
LIMO: Less is More for Reasoning
von: Ye, Yixin, et al.
Veröffentlicht: (2025) -
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024) -
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
von: Zhuang, Xialie, et al.
Veröffentlicht: (2025)