RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Junhao, Huang, Wenrui, Wang, Weidong, Li, Zhenwen, Hu, Tiancheng, Liu, Zhixia, Chen, Xusheng, Xie, Tao, Shan, Yizhou |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A methodological framework for Resilience as a Service (RaaS) in multimodal urban transportation networks
par: Jaber, Sara, et autres
Publié: (2024)
par: Jaber, Sara, et autres
Publié: (2024)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
par: Hu, Junhao, et autres
Publié: (2024)
par: Hu, Junhao, et autres
Publié: (2024)
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
par: Su, Tiancheng, et autres
Publié: (2025)
par: Su, Tiancheng, et autres
Publié: (2025)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
par: Hu, Junhao, et autres
Publié: (2026)
par: Hu, Junhao, et autres
Publié: (2026)
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
par: Liang, Zhenwen, et autres
Publié: (2024)
par: Liang, Zhenwen, et autres
Publié: (2024)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
par: Wan, Guangya, et autres
Publié: (2024)
par: Wan, Guangya, et autres
Publié: (2024)
Graph-Augmented Reasoning: Evolving Step-by-Step Knowledge Graph Retrieval for LLM Reasoning
par: Wu, Wenjie, et autres
Publié: (2025)
par: Wu, Wenjie, et autres
Publié: (2025)
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
par: Huang, Chen, et autres
Publié: (2025)
par: Huang, Chen, et autres
Publié: (2025)
Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoning
par: Shan, Lianlei, et autres
Publié: (2026)
par: Shan, Lianlei, et autres
Publié: (2026)
Empowering LLM Agents with Geospatial Awareness: Toward Grounded Reasoning for Wildfire Response
par: Chen, Yiheng, et autres
Publié: (2025)
par: Chen, Yiheng, et autres
Publié: (2025)
Scalable and Accurate Graph Reasoning with LLM-based Multi-Agents
par: Hu, Yuwei, et autres
Publié: (2024)
par: Hu, Yuwei, et autres
Publié: (2024)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
par: Panaganti, Kishan, et autres
Publié: (2026)
par: Panaganti, Kishan, et autres
Publié: (2026)
Scaling Attention via Feature Sparsity
par: Xie, Yan, et autres
Publié: (2026)
par: Xie, Yan, et autres
Publié: (2026)
Dynamic Parallel Tree Search for Efficient LLM Reasoning
par: Ding, Yifu, et autres
Publié: (2025)
par: Ding, Yifu, et autres
Publié: (2025)
Scaling-Aware Adapter for Structure-Grounded LLM Reasoning
par: Jing, Zihao, et autres
Publié: (2026)
par: Jing, Zihao, et autres
Publié: (2026)
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
par: Yang, Tianyu, et autres
Publié: (2026)
par: Yang, Tianyu, et autres
Publié: (2026)
Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness
par: Chen, Zizhao, et autres
Publié: (2026)
par: Chen, Zizhao, et autres
Publié: (2026)
Enhanced Cyber Threat Intelligence by Network Forensic Analysis for Ransomware as a Service(RaaS) Malwares
par: P, Sharmila S
Publié: (2026)
par: P, Sharmila S
Publié: (2026)
FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning
par: Luo, Haozheng, et autres
Publié: (2026)
par: Luo, Haozheng, et autres
Publié: (2026)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
par: Fan, Wang, et autres
Publié: (2026)
par: Fan, Wang, et autres
Publié: (2026)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
par: Chen, Yi, et autres
Publié: (2025)
par: Chen, Yi, et autres
Publié: (2025)
Skill-Conditioned Gated Self-Distillation for LLM Reasoning
par: Huang, Jiazhen, et autres
Publié: (2026)
par: Huang, Jiazhen, et autres
Publié: (2026)
SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs
par: Xu, Jiaming, et autres
Publié: (2025)
par: Xu, Jiaming, et autres
Publié: (2025)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
par: Liu, Yifei, et autres
Publié: (2025)
par: Liu, Yifei, et autres
Publié: (2025)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
par: Liang, Zhenwen, et autres
Publié: (2025)
par: Liang, Zhenwen, et autres
Publié: (2025)
RelayLLM: Efficient Reasoning via Collaborative Decoding
par: Huang, Chengsong, et autres
Publié: (2026)
par: Huang, Chengsong, et autres
Publié: (2026)
Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models
par: Huang, Yizhou, et autres
Publié: (2024)
par: Huang, Yizhou, et autres
Publié: (2024)
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs
par: Shu, Yubo, et autres
Publié: (2025)
par: Shu, Yubo, et autres
Publié: (2025)
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
par: Jiang, Xinyan, et autres
Publié: (2026)
par: Jiang, Xinyan, et autres
Publié: (2026)
Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
par: Gao, Xianjun, et autres
Publié: (2025)
par: Gao, Xianjun, et autres
Publié: (2025)
IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
par: Shen, Yuanzhe, et autres
Publié: (2025)
par: Shen, Yuanzhe, et autres
Publié: (2025)
Mixture of Reasonings: Teach Large Language Models to Reason with Adaptive Strategies
par: Xiong, Tao, et autres
Publié: (2025)
par: Xiong, Tao, et autres
Publié: (2025)
Token-Budget-Aware LLM Reasoning
par: Han, Tingxu, et autres
Publié: (2024)
par: Han, Tingxu, et autres
Publié: (2024)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
par: Qin, Jialong, et autres
Publié: (2025)
par: Qin, Jialong, et autres
Publié: (2025)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
par: Chen, Haoyang, et autres
Publié: (2026)
par: Chen, Haoyang, et autres
Publié: (2026)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
par: Lu, Jinghui, et autres
Publié: (2025)
par: Lu, Jinghui, et autres
Publié: (2025)
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
par: Huang, Zixuan, et autres
Publié: (2026)
par: Huang, Zixuan, et autres
Publié: (2026)
TheraAgent: Multi-Agent Framework with Self-Evolving Memory and Evidence-Calibrated Reasoning for PET Theranostics
par: Chen, Zhihao, et autres
Publié: (2026)
par: Chen, Zhihao, et autres
Publié: (2026)
Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
par: Cai, Wenrui, et autres
Publié: (2025)
par: Cai, Wenrui, et autres
Publié: (2025)
Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning
par: Guo, Chenyou, et autres
Publié: (2026)
par: Guo, Chenyou, et autres
Publié: (2026)
Documents similaires
-
A methodological framework for Resilience as a Service (RaaS) in multimodal urban transportation networks
par: Jaber, Sara, et autres
Publié: (2024) -
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
par: Hu, Junhao, et autres
Publié: (2024) -
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
par: Su, Tiancheng, et autres
Publié: (2025) -
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
par: Hu, Junhao, et autres
Publié: (2026) -
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
par: Liang, Zhenwen, et autres
Publié: (2024)