Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fei, Weizhi, Niu, Xueyan, Xie, Guoqing, Liu, Yingqing, Bai, Bo, Han, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Retrieval Meets Reasoning: Dynamic In-Context Editing for Long-Text Understanding
von: Fei, Weizhi, et al.
Veröffentlicht: (2024)
von: Fei, Weizhi, et al.
Veröffentlicht: (2024)
Characterizing Prompt Compression Methods for Long Context Inference
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
NeuralDB: Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
von: Ge, Suyu, et al.
Veröffentlicht: (2024)
von: Ge, Suyu, et al.
Veröffentlicht: (2024)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
CCF: A Context Compression Framework for Efficient Long-Sequence Language Modeling
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
von: Li, Wenhao, et al.
Veröffentlicht: (2025)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
von: Liu, Siran, et al.
Veröffentlicht: (2026)
von: Liu, Siran, et al.
Veröffentlicht: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
von: Tian, Han, et al.
Veröffentlicht: (2026)
von: Tian, Han, et al.
Veröffentlicht: (2026)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
von: Lu, Yi, et al.
Veröffentlicht: (2024)
von: Lu, Yi, et al.
Veröffentlicht: (2024)
Fovea Transformer: Efficient Long-Context Modeling with Structured Fine-to-Coarse Attention
von: He, Ziwei, et al.
Veröffentlicht: (2023)
von: He, Ziwei, et al.
Veröffentlicht: (2023)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2023)
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2023)
Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios
von: Tang, Jiwei, et al.
Veröffentlicht: (2024)
von: Tang, Jiwei, et al.
Veröffentlicht: (2024)
Efficient Dynamic Clustering-Based Document Compression for Retrieval-Augmented-Generation
von: Li, Weitao, et al.
Veröffentlicht: (2025)
von: Li, Weitao, et al.
Veröffentlicht: (2025)
ICPC: In-context Prompt Compression with Faster Inference
von: Yu, Ziyang, et al.
Veröffentlicht: (2025)
von: Yu, Ziyang, et al.
Veröffentlicht: (2025)
Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
von: Liskavets, Barys, et al.
Veröffentlicht: (2024)
von: Liskavets, Barys, et al.
Veröffentlicht: (2024)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
von: Mao, Weian, et al.
Veröffentlicht: (2026)
von: Mao, Weian, et al.
Veröffentlicht: (2026)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
von: Ma, Da, et al.
Veröffentlicht: (2024)
von: Ma, Da, et al.
Veröffentlicht: (2024)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
von: Peng, Junjie, et al.
Veröffentlicht: (2026)
von: Peng, Junjie, et al.
Veröffentlicht: (2026)
Evaluating Zero-Shot Long-Context LLM Compression
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
von: Song, Woomin, et al.
Veröffentlicht: (2025)
von: Song, Woomin, et al.
Veröffentlicht: (2025)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
Latent-Condensed Transformer for Efficient Long Context Modeling
von: You, Zeng, et al.
Veröffentlicht: (2026)
von: You, Zeng, et al.
Veröffentlicht: (2026)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
von: Wan, Luanbo, et al.
Veröffentlicht: (2025)
von: Wan, Luanbo, et al.
Veröffentlicht: (2025)
Inference Scaling for Long-Context Retrieval Augmented Generation
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
ATACompressor: Adaptive Task-Aware Compression for Efficient Long-Context Processing in LLMs
von: Li, Xuancheng, et al.
Veröffentlicht: (2026)
von: Li, Xuancheng, et al.
Veröffentlicht: (2026)
Long Context Compression with Activation Beacon
von: Zhang, Peitian, et al.
Veröffentlicht: (2024)
von: Zhang, Peitian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Retrieval Meets Reasoning: Dynamic In-Context Editing for Long-Text Understanding
von: Fei, Weizhi, et al.
Veröffentlicht: (2024) -
Characterizing Prompt Compression Methods for Long Context Inference
von: Jha, Siddharth, et al.
Veröffentlicht: (2024) -
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024) -
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025) -
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)