PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Lizhe, Zhou, Binjia, Ge, Yuyao, Chen, Jiayi, NI, Shiguang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
von: Figliolia, Tomas, et al.
Veröffentlicht: (2025)
von: Figliolia, Tomas, et al.
Veröffentlicht: (2025)
EFPC: Towards Efficient and Flexible Prompt Compression
von: Cao, Yun-Hao, et al.
Veröffentlicht: (2025)
von: Cao, Yun-Hao, et al.
Veröffentlicht: (2025)
Efficient and Effective Prompt Tuning via Prompt Decomposition and Compressed Outer Product
von: Lan, Pengxiang, et al.
Veröffentlicht: (2025)
von: Lan, Pengxiang, et al.
Veröffentlicht: (2025)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation
von: Lu, Yao, et al.
Veröffentlicht: (2023)
von: Lu, Yao, et al.
Veröffentlicht: (2023)
Recurrent Context Compression: Efficiently Expanding the Context Window of LLM
von: Huang, Chensen, et al.
Veröffentlicht: (2024)
von: Huang, Chensen, et al.
Veröffentlicht: (2024)
ACCEPT: Adaptive Codebook for Composite and Efficient Prompt Tuning
von: Lin, Yu-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Yu-Chen, et al.
Veröffentlicht: (2024)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
von: Leemann, Tobias, et al.
Veröffentlicht: (2024)
von: Leemann, Tobias, et al.
Veröffentlicht: (2024)
Asking LLMs to Verify First is Almost Free Lunch
von: Wu, Shiguang, et al.
Veröffentlicht: (2025)
von: Wu, Shiguang, et al.
Veröffentlicht: (2025)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
von: Gim, In, et al.
Veröffentlicht: (2023)
von: Gim, In, et al.
Veröffentlicht: (2023)
Discrete Prompt Compression with Reinforcement Learning
von: Jung, Hoyoun, et al.
Veröffentlicht: (2023)
von: Jung, Hoyoun, et al.
Veröffentlicht: (2023)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
Efficient Attention Mechanisms for Large Language Models: A Survey
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspective
von: Peng, Dengyun, et al.
Veröffentlicht: (2025)
von: Peng, Dengyun, et al.
Veröffentlicht: (2025)
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2026)
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2026)
Language Models Resist Alignment: Evidence From Data Compression
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
ICPC: In-context Prompt Compression with Faster Inference
von: Yu, Ziyang, et al.
Veröffentlicht: (2025)
von: Yu, Ziyang, et al.
Veröffentlicht: (2025)
Parse Trees Guided LLM Prompt Compression
von: Mao, Wenhao, et al.
Veröffentlicht: (2024)
von: Mao, Wenhao, et al.
Veröffentlicht: (2024)
Learning to Compress Prompt in Natural Language Formats
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
von: Tang, Yaohua, et al.
Veröffentlicht: (2025)
von: Tang, Yaohua, et al.
Veröffentlicht: (2025)
Importance of Prompt Optimisation for Error Detection in Medical Notes Using Language Models
von: Myles, Craig, et al.
Veröffentlicht: (2026)
von: Myles, Craig, et al.
Veröffentlicht: (2026)
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization
von: Sahu, Gaurav, et al.
Veröffentlicht: (2023)
von: Sahu, Gaurav, et al.
Veröffentlicht: (2023)
Spectral Attention Steering for Prompt Highlighting
von: Li, Weixian Waylon, et al.
Veröffentlicht: (2026)
von: Li, Weixian Waylon, et al.
Veröffentlicht: (2026)
Efficient Streaming Language Models with Attention Sinks
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
von: Yang, Ge, et al.
Veröffentlicht: (2024)
von: Yang, Ge, et al.
Veröffentlicht: (2024)
ConsPrompt: Exploiting Contrastive Samples for Fewshot Prompt Learning
von: Weng, Jinta, et al.
Veröffentlicht: (2022)
von: Weng, Jinta, et al.
Veröffentlicht: (2022)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
Adapting LLMs for Efficient Context Processing through Soft Prompt Compression
von: Wang, Cangqing, et al.
Veröffentlicht: (2024)
von: Wang, Cangqing, et al.
Veröffentlicht: (2024)
SCOPE: A Generative Approach for LLM Prompt Compression
von: Zhang, Tinghui, et al.
Veröffentlicht: (2025)
von: Zhang, Tinghui, et al.
Veröffentlicht: (2025)
An Empirical Study on Prompt Compression for Large Language Models
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
URL: Universal Referential Knowledge Linking via Task-instructed Representation Compression
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
von: Lee, Gisang, et al.
Veröffentlicht: (2024)
von: Lee, Gisang, et al.
Veröffentlicht: (2024)
Attention Mechanism and Context Modeling System for Text Mining Machine Translation
von: Zhang, Yuwei, et al.
Veröffentlicht: (2024)
von: Zhang, Yuwei, et al.
Veröffentlicht: (2024)
$\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
von: Xu, Yue, et al.
Veröffentlicht: (2024)
von: Xu, Yue, et al.
Veröffentlicht: (2024)
Optimizing Length Compression in Large Reasoning Models
von: Cheng, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Cheng, Zhengxiang, et al.
Veröffentlicht: (2025)
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
von: Liu, Chi, et al.
Veröffentlicht: (2026)
von: Liu, Chi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
von: Figliolia, Tomas, et al.
Veröffentlicht: (2025) -
EFPC: Towards Efficient and Flexible Prompt Compression
von: Cao, Yun-Hao, et al.
Veröffentlicht: (2025) -
Efficient and Effective Prompt Tuning via Prompt Decomposition and Compressed Outer Product
von: Lan, Pengxiang, et al.
Veröffentlicht: (2025) -
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
von: Hu, Jinwu, et al.
Veröffentlicht: (2025) -
Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation
von: Lu, Yao, et al.
Veröffentlicht: (2023)