LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Zhuoshi, Wu, Qianhui, Jiang, Huiqiang, Xia, Menglin, Luo, Xufang, Zhang, Jue, Lin, Qingwei, Rühle, Victor, Yang, Yuqing, Lin, Chin-Yew, Zhao, H. Vicky, Qiu, Lili, Zhang, Dongmei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
by: Jiang, Huiqiang, et al.
Published: (2023)
by: Jiang, Huiqiang, et al.
Published: (2023)
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
by: Shandilya, Shivam, et al.
Published: (2024)
by: Shandilya, Shivam, et al.
Published: (2024)
On Memory Construction and Retrieval for Personalized Conversational Agents
by: Pan, Zhuoshi, et al.
Published: (2025)
by: Pan, Zhuoshi, et al.
Published: (2025)
Mitigate Position Bias in Large Language Models via Scaling a Single Dimension
by: Yu, Yijiong, et al.
Published: (2024)
by: Yu, Yijiong, et al.
Published: (2024)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
by: Zhang, Jue, et al.
Published: (2025)
by: Zhang, Jue, et al.
Published: (2025)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
by: Jiang, Huiqiang, et al.
Published: (2024)
by: Jiang, Huiqiang, et al.
Published: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
by: Li, Yucheng, et al.
Published: (2025)
by: Li, Yucheng, et al.
Published: (2025)
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
by: Chung, Tsz Ting, et al.
Published: (2024)
by: Chung, Tsz Ting, et al.
Published: (2024)
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
by: Huang, Sterling, et al.
Published: (2026)
by: Huang, Sterling, et al.
Published: (2026)
MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
by: Li, Yucheng, et al.
Published: (2025)
by: Li, Yucheng, et al.
Published: (2025)
SCBench: A KV Cache-Centric Analysis of Long-Context Methods
by: Li, Yucheng, et al.
Published: (2024)
by: Li, Yucheng, et al.
Published: (2024)
A Tale of Two Graphs: Separating Knowledge Exploration from Outline Structure for Open-Ended Deep Research
by: Shi, Zhuofan, et al.
Published: (2026)
by: Shi, Zhuofan, et al.
Published: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
by: Tan, Rongyuan, et al.
Published: (2026)
by: Tan, Rongyuan, et al.
Published: (2026)
Navigating the Unknown: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks
by: Peng, Yingzhe, et al.
Published: (2024)
by: Peng, Yingzhe, et al.
Published: (2024)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
by: Qian, Jiaxu, et al.
Published: (2025)
by: Qian, Jiaxu, et al.
Published: (2025)
Accelerating Prefilling via Decoding-time Contribution Sparsity
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
Position Engineering: Boosting Large Language Models through Positional Information Manipulation
by: He, Zhiyuan, et al.
Published: (2024)
by: He, Zhiyuan, et al.
Published: (2024)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
by: Hu, Lingxiang, et al.
Published: (2025)
by: Hu, Lingxiang, et al.
Published: (2025)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
by: Zhang, Yike, et al.
Published: (2025)
by: Zhang, Yike, et al.
Published: (2025)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
by: Wang, Shouju, et al.
Published: (2025)
by: Wang, Shouju, et al.
Published: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
by: Fu, Wenjie, et al.
Published: (2026)
by: Fu, Wenjie, et al.
Published: (2026)
Hybrid-RACA: Hybrid Retrieval-Augmented Composition Assistance for Real-time Text Prediction
by: Xia, Menglin, et al.
Published: (2023)
by: Xia, Menglin, et al.
Published: (2023)
P‐3.9: The Influence of Parallax and Shape Type Factors on the Perception of AR Equipment in Dark Environment
by: Huiqiang Xia, et al.
Published: (2024)
by: Huiqiang Xia, et al.
Published: (2024)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
by: Zhang, Zhiyang, et al.
Published: (2024)
by: Zhang, Zhiyang, et al.
Published: (2024)
Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
by: Cheng, Sitao, et al.
Published: (2024)
by: Cheng, Sitao, et al.
Published: (2024)
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
by: Liao, Mengqi, et al.
Published: (2026)
by: Liao, Mengqi, et al.
Published: (2026)
AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based Agents
by: Lu, Junting, et al.
Published: (2024)
by: Lu, Junting, et al.
Published: (2024)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
by: Fu, Jia, et al.
Published: (2024)
by: Fu, Jia, et al.
Published: (2024)
Sharingan: Extract User Action Sequence from Desktop Recordings
by: Chen, Yanting, et al.
Published: (2024)
by: Chen, Yanting, et al.
Published: (2024)
From Engagement to Incorporation: What Goes Beyond Online Peer Interaction in L2 Writing
by: Dongmei Zhang, et al.
Published: (2025)
by: Dongmei Zhang, et al.
Published: (2025)
Revisiting Auxiliary Losses for Conditional Depth Routing: An Empirical Study
by: Lin, Qingwei
Published: (2026)
by: Lin, Qingwei
Published: (2026)
Enabling Autonomic Microservice Management through Self-Learning Agents
by: Yu, Fenglin, et al.
Published: (2025)
by: Yu, Fenglin, et al.
Published: (2025)
Designing Network Algorithms via Large Language Models
by: He, Zhiyuan, et al.
Published: (2024)
by: He, Zhiyuan, et al.
Published: (2024)
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
by: Zhang, Yudi, et al.
Published: (2025)
by: Zhang, Yudi, et al.
Published: (2025)
Similar Items
-
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
by: Jiang, Huiqiang, et al.
Published: (2023) -
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
by: Shandilya, Shivam, et al.
Published: (2024) -
On Memory Construction and Retrieval for Personalized Conversational Agents
by: Pan, Zhuoshi, et al.
Published: (2025) -
Mitigate Position Bias in Large Language Models via Scaling a Single Dimension
by: Yu, Yijiong, et al.
Published: (2024) -
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
by: Zhang, Jue, et al.
Published: (2025)