Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Liskavets, Barys, Ushakov, Maxim, Roy, Shuvendu, Klibanov, Mark, Etemad, Ali, Luke, Shane |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor
by: Liskavets, Barys, et al.
Published: (2025)
by: Liskavets, Barys, et al.
Published: (2025)
Scaling Up Semi-supervised Learning with Unconstrained Unlabelled Data
by: Roy, Shuvendu, et al.
Published: (2023)
by: Roy, Shuvendu, et al.
Published: (2023)
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
by: Roy, Shuvendu, et al.
Published: (2025)
by: Roy, Shuvendu, et al.
Published: (2025)
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
A Bag of Tricks for Few-Shot Class-Incremental Learning
by: Roy, Shuvendu, et al.
Published: (2024)
by: Roy, Shuvendu, et al.
Published: (2024)
Consistency-guided Prompt Learning for Vision-Language Models
by: Roy, Shuvendu, et al.
Published: (2023)
by: Roy, Shuvendu, et al.
Published: (2023)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
Evaluating LLM-driven User-Intent Formalization for Verification-Aware Languages
by: Lahiri, Shuvendu K.
Published: (2024)
by: Lahiri, Shuvendu K.
Published: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
Tiny Transformers Excel at Sentence Compression
by: Belcak, Peter, et al.
Published: (2024)
by: Belcak, Peter, et al.
Published: (2024)
Speech Emotion Recognition with Distilled Prosodic and Linguistic Affect Representations
by: Shome, Debaditya, et al.
Published: (2023)
by: Shome, Debaditya, et al.
Published: (2023)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
Prompt-SAW: Leveraging Relation-Aware Graphs for Textual Prompt Compression
by: Ali, Muhammad Asif, et al.
Published: (2024)
by: Ali, Muhammad Asif, et al.
Published: (2024)
Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression
by: Trukhina, Natalia, et al.
Published: (2026)
by: Trukhina, Natalia, et al.
Published: (2026)
LLM In-Context Recall is Prompt Dependent
by: Machlab, Daniel, et al.
Published: (2024)
by: Machlab, Daniel, et al.
Published: (2024)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
by: Behnam, Payman, et al.
Published: (2025)
by: Behnam, Payman, et al.
Published: (2025)
Impact of Strategic Sampling and Supervision Policies on Semi-supervised Learning
by: Roy, Shuvendu, et al.
Published: (2022)
by: Roy, Shuvendu, et al.
Published: (2022)
Exploring the Boundaries of Semi-Supervised Facial Expression Recognition using In-Distribution, Out-of-Distribution, and Unconstrained Data
by: Roy, Shuvendu, et al.
Published: (2023)
by: Roy, Shuvendu, et al.
Published: (2023)
Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data
by: de Campos, Andresa Rodrigues, et al.
Published: (2026)
by: de Campos, Andresa Rodrigues, et al.
Published: (2026)
ProCut: LLM Prompt Compression via Attribution Estimation
by: Xu, Zhentao, et al.
Published: (2025)
by: Xu, Zhentao, et al.
Published: (2025)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
by: Shandilya, Shivam, et al.
Published: (2024)
by: Shandilya, Shivam, et al.
Published: (2024)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
by: Jiang, Huiqiang, et al.
Published: (2023)
by: Jiang, Huiqiang, et al.
Published: (2023)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
by: Sunesh, Aman, et al.
Published: (2026)
by: Sunesh, Aman, et al.
Published: (2026)
Communication Compression for Tensor Parallel LLM Inference
by: Hansen-Palmus, Jan, et al.
Published: (2024)
by: Hansen-Palmus, Jan, et al.
Published: (2024)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
by: Adamska, Marta, et al.
Published: (2025)
by: Adamska, Marta, et al.
Published: (2025)
Adapting LLMs for Efficient Context Processing through Soft Prompt Compression
by: Wang, Cangqing, et al.
Published: (2024)
by: Wang, Cangqing, et al.
Published: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026)
by: Kim, Jang-Hyun, et al.
Published: (2026)
Calibrate to Discriminate: Improve In-Context Learning with Label-Free Comparative Inference
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
ChameleonLLM: Batch-Aware Dynamic Low-Rank Adaptation via Inference-Time Clusters
by: Yuksel, Kamer Ali, et al.
Published: (2025)
by: Yuksel, Kamer Ali, et al.
Published: (2025)
Improving the Efficiency of Long Document Classification using Sentence Ranking Approach
by: Kokate, Prathamesh, et al.
Published: (2025)
by: Kokate, Prathamesh, et al.
Published: (2025)
Improving Sampling Methods for Fine-tuning SentenceBERT in Text Streams
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
Enhancing Foundation Models in Transaction Understanding with LLM-based Sentence Embeddings
by: Fan, Xiran, et al.
Published: (2025)
by: Fan, Xiran, et al.
Published: (2025)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
Efficient Zero-Shot Long Document Classification by Reducing Context Through Sentence Ranking
by: Kokate, Prathamesh, et al.
Published: (2025)
by: Kokate, Prathamesh, et al.
Published: (2025)
Similar Items
-
Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor
by: Liskavets, Barys, et al.
Published: (2025) -
Scaling Up Semi-supervised Learning with Unconstrained Unlabelled Data
by: Roy, Shuvendu, et al.
Published: (2023) -
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
by: Roy, Shuvendu, et al.
Published: (2025) -
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024) -
A Bag of Tricks for Few-Shot Class-Incremental Learning
by: Roy, Shuvendu, et al.
Published: (2024)