DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yi, Li, Zuchao, Zhao, Hai, Qi, Baoyuan, Liu, Guoming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
by: Shi, Luohe, et al.
Published: (2025)
by: Shi, Luohe, et al.
Published: (2025)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
by: Yang, Haoqi, et al.
Published: (2025)
by: Yang, Haoqi, et al.
Published: (2025)
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
by: Shi, Luohe, et al.
Published: (2025)
by: Shi, Luohe, et al.
Published: (2025)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
by: Hu, Jiliang, et al.
Published: (2025)
by: Hu, Jiliang, et al.
Published: (2025)
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)
by: Yao, Yao, et al.
Published: (2023)
SirLLM: Streaming Infinite Retentive LLM
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
by: Pan, Zhuoshi, et al.
Published: (2024)
by: Pan, Zhuoshi, et al.
Published: (2024)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
by: Shi, Luohe, et al.
Published: (2024)
by: Shi, Luohe, et al.
Published: (2024)
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
by: Shi, Luohe, et al.
Published: (2024)
by: Shi, Luohe, et al.
Published: (2024)
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
by: Chung, Tsz Ting, et al.
Published: (2024)
by: Chung, Tsz Ting, et al.
Published: (2024)
Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor
by: Liskavets, Barys, et al.
Published: (2025)
by: Liskavets, Barys, et al.
Published: (2025)
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba
by: Zou, Yuchen, et al.
Published: (2024)
by: Zou, Yuchen, et al.
Published: (2024)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
by: Zeng, Xiangke, et al.
Published: (2024)
by: Zeng, Xiangke, et al.
Published: (2024)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
by: Zhang, Zihong, et al.
Published: (2026)
by: Zhang, Zihong, et al.
Published: (2026)
Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios
by: Tang, Jiwei, et al.
Published: (2024)
by: Tang, Jiwei, et al.
Published: (2024)
Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations
by: Chen, Nuo, et al.
Published: (2024)
by: Chen, Nuo, et al.
Published: (2024)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
by: Song, Weixi, et al.
Published: (2023)
by: Song, Weixi, et al.
Published: (2023)
Semantics-Preserved Distortion for Personal Privacy Protection in Information Management
by: Li, Jiajia, et al.
Published: (2022)
by: Li, Jiajia, et al.
Published: (2022)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
by: Zhang, Zihong, et al.
Published: (2025)
by: Zhang, Zihong, et al.
Published: (2025)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
by: Yao, Yao, et al.
Published: (2025)
by: Yao, Yao, et al.
Published: (2025)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
by: Hu, Jiliang, et al.
Published: (2024)
by: Hu, Jiliang, et al.
Published: (2024)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
by: He, Xingyang, et al.
Published: (2025)
by: He, Xingyang, et al.
Published: (2025)
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
EFPC: Towards Efficient and Flexible Prompt Compression
by: Cao, Yun-Hao, et al.
Published: (2025)
by: Cao, Yun-Hao, et al.
Published: (2025)
Self-supervised Attribute-aware Dynamic Preference Ranking Alignment
by: Yang, Hongyu, et al.
Published: (2025)
by: Yang, Hongyu, et al.
Published: (2025)
Transferable Modeling Strategies for Low-Resource LLM Tasks: A Prompt and Alignment-Based Approach
by: Lyu, Shuangquan, et al.
Published: (2025)
by: Lyu, Shuangquan, et al.
Published: (2025)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
by: Guo, Jiani, et al.
Published: (2025)
by: Guo, Jiani, et al.
Published: (2025)
FamiCom: Further Demystifying Prompts for Language Models with Task-Agnostic Performance Estimation
by: Li, Bangzheng, et al.
Published: (2024)
by: Li, Bangzheng, et al.
Published: (2024)
Density-aware Soft Context Compression with Semi-Dynamic Compression Ratio
by: Yu, Yijiong, et al.
Published: (2026)
by: Yu, Yijiong, et al.
Published: (2026)
Style-Compress: An LLM-Based Prompt Compression Framework Considering Task-Specific Styles
by: Pu, Xiao, et al.
Published: (2024)
by: Pu, Xiao, et al.
Published: (2024)
Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning
by: Yang, Juncheng, et al.
Published: (2024)
by: Yang, Juncheng, et al.
Published: (2024)
TAPO: Task-Referenced Adaptation for Prompt Optimization
by: Luo, Wenxin, et al.
Published: (2025)
by: Luo, Wenxin, et al.
Published: (2025)
A Novel Multi-Stage Prompting Approach for Language Agnostic MCQ Generation using GPT
by: Maity, Subhankar, et al.
Published: (2024)
by: Maity, Subhankar, et al.
Published: (2024)
Unveiling Vulnerability of Self-Attention
by: Liong, Khai Jiet, et al.
Published: (2024)
by: Liong, Khai Jiet, et al.
Published: (2024)
Similar Items
-
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
by: Shi, Luohe, et al.
Published: (2025) -
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
by: Yang, Haoqi, et al.
Published: (2025) -
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
by: Shi, Luohe, et al.
Published: (2025) -
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
by: Hu, Jiliang, et al.
Published: (2025) -
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
by: Zhao, Yi, et al.
Published: (2025)