Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Yaohua, Hu, Zhicheng, Cheng, Kun, Mo, Fan, Lv, Qiheng, Wang, Hua, Chen, Zhi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
di: Wu, Wenhao, et al.
Pubblicazione: (2026)
di: Wu, Wenhao, et al.
Pubblicazione: (2026)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
di: Zhang, Hanzhi, et al.
Pubblicazione: (2025)
di: Zhang, Hanzhi, et al.
Pubblicazione: (2025)
PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator
di: Kong, Chuyi, et al.
Pubblicazione: (2023)
di: Kong, Chuyi, et al.
Pubblicazione: (2023)
Self-Selected Attention Span for Accelerating Large Language Model Inference
di: Jin, Tian, et al.
Pubblicazione: (2024)
di: Jin, Tian, et al.
Pubblicazione: (2024)
MRJ-Agent: An Effective Jailbreak Agent for Multi-Round Dialogue
di: Wang, Fengxiang, et al.
Pubblicazione: (2024)
di: Wang, Fengxiang, et al.
Pubblicazione: (2024)
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
di: Guan, Zhong, et al.
Pubblicazione: (2025)
di: Guan, Zhong, et al.
Pubblicazione: (2025)
SEKI: Self-Evolution and Knowledge Inspiration based Neural Architecture Search via Large Language Models
di: Cai, Zicheng, et al.
Pubblicazione: (2025)
di: Cai, Zicheng, et al.
Pubblicazione: (2025)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
di: Cheng, Yixin, et al.
Pubblicazione: (2024)
di: Cheng, Yixin, et al.
Pubblicazione: (2024)
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
di: Ouyang, Haojie, et al.
Pubblicazione: (2025)
di: Ouyang, Haojie, et al.
Pubblicazione: (2025)
FlashEVA: Accelerating LLM inference via Efficient Attention
di: Kostelec, Juan Gabriel, et al.
Pubblicazione: (2025)
di: Kostelec, Juan Gabriel, et al.
Pubblicazione: (2025)
Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue
di: Wang, Jian, et al.
Pubblicazione: (2024)
di: Wang, Jian, et al.
Pubblicazione: (2024)
PET-SQL: A Prompt-Enhanced Two-Round Refinement of Text-to-SQL with Cross-consistency
di: Li, Zhishuai, et al.
Pubblicazione: (2024)
di: Li, Zhishuai, et al.
Pubblicazione: (2024)
Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
di: Liu, Kai, et al.
Pubblicazione: (2025)
di: Liu, Kai, et al.
Pubblicazione: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
di: Skorobogat, Ronald, et al.
Pubblicazione: (2026)
di: Skorobogat, Ronald, et al.
Pubblicazione: (2026)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning
di: Attia, Ahmed, et al.
Pubblicazione: (2026)
di: Attia, Ahmed, et al.
Pubblicazione: (2026)
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
di: Yung, Canaan, et al.
Pubblicazione: (2024)
di: Yung, Canaan, et al.
Pubblicazione: (2024)
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
di: Cheng, Wenhua, et al.
Pubblicazione: (2025)
di: Cheng, Wenhua, et al.
Pubblicazione: (2025)
Inference-Friendly Models With MixAttention
di: Rajput, Shashank, et al.
Pubblicazione: (2024)
di: Rajput, Shashank, et al.
Pubblicazione: (2024)
Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents
di: Ellawela, Suveen
Pubblicazione: (2026)
di: Ellawela, Suveen
Pubblicazione: (2026)
Star Attention: Efficient LLM Inference over Long Sequences
di: Acharya, Shantanu, et al.
Pubblicazione: (2024)
di: Acharya, Shantanu, et al.
Pubblicazione: (2024)
ReAttn: Improving Attention-based Re-ranking via Attention Re-weighting
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
di: Tian, Yuxing, et al.
Pubblicazione: (2026)
HSR-Enhanced Sparse Attention Acceleration
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
di: Hua, Ermo, et al.
Pubblicazione: (2024)
di: Hua, Ermo, et al.
Pubblicazione: (2024)
Mixture-of-Depths Attention
di: Zhu, Lianghui, et al.
Pubblicazione: (2026)
di: Zhu, Lianghui, et al.
Pubblicazione: (2026)
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
di: Siska, Charlotte, et al.
Pubblicazione: (2025)
di: Siska, Charlotte, et al.
Pubblicazione: (2025)
OUNLP at TSAR 2025 Shared Task: Multi-Round Text Simplifier via Code Generation
di: Huynh, Cuong, et al.
Pubblicazione: (2025)
di: Huynh, Cuong, et al.
Pubblicazione: (2025)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
di: Gim, In, et al.
Pubblicazione: (2023)
di: Gim, In, et al.
Pubblicazione: (2023)
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
di: Tong, Qinyue, et al.
Pubblicazione: (2025)
di: Tong, Qinyue, et al.
Pubblicazione: (2025)
IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues
di: Yang, Diji, et al.
Pubblicazione: (2024)
di: Yang, Diji, et al.
Pubblicazione: (2024)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
A Framework for Inference Inspired by Human Memory Mechanisms
di: Zeng, Xiangyu, et al.
Pubblicazione: (2023)
di: Zeng, Xiangyu, et al.
Pubblicazione: (2023)
PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead
di: Tan, Tao, et al.
Pubblicazione: (2024)
di: Tan, Tao, et al.
Pubblicazione: (2024)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
di: Lin, Gang, et al.
Pubblicazione: (2026)
di: Lin, Gang, et al.
Pubblicazione: (2026)
ReAttention: Training-Free Infinite Context with Finite Attention Scope
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
di: Chen, Lizhe, et al.
Pubblicazione: (2025)
di: Chen, Lizhe, et al.
Pubblicazione: (2025)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
di: Zhu, Qianchao, et al.
Pubblicazione: (2024) -
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
di: Wu, Wenhao, et al.
Pubblicazione: (2026) -
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
di: Cheng, Wenhua, et al.
Pubblicazione: (2023) -
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
di: Zhang, Hanzhi, et al.
Pubblicazione: (2025) -
PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator
di: Kong, Chuyi, et al.
Pubblicazione: (2023)