PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Runsong, Liu, Shilei, Tang, Jiwei, Liu, Langming, Chen, Haibin, Zhang, Weidong, Yuan, Yujin, Xiao, Tong, Zhu, Jingbo, Su, Wenbo, Zheng, Bo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915877413191680
author Zhao, Runsong
Liu, Shilei
Tang, Jiwei
Liu, Langming
Chen, Haibin
Zhang, Weidong
Yuan, Yujin
Xiao, Tong
Zhu, Jingbo
Su, Wenbo
Zheng, Bo
author_facet Zhao, Runsong
Liu, Shilei
Tang, Jiwei
Liu, Langming
Chen, Haibin
Zhang, Weidong
Yuan, Yujin
Xiao, Tong
Zhu, Jingbo
Su, Wenbo
Zheng, Bo
contents While context compression can mitigate the growing inference costs of Large Language Models (LLMs) by shortening contexts, existing methods that specify a target compression ratio or length suffer from unpredictable performance degradation, hindering their reliable deployment. We introduce a paradigm shift to Performance-oriented Context Compression (PoC), where developers specify an acceptable performance floor instead of a compression ratio. PoC employs a lightweight performance predictor to automatically find the most aggressive compression ratio that satisfies this constraint before steering an off-the-shelf compressor. We design and compare two predictor variants: a simple context-agnostic predictor and a more sophisticated context-aware one that considers the input's inherent compressibility. On both question-answering and summarization benchmarks, the context-aware predictor consistently achieves lower performance prediction error than the context-agnostic predictor, while the resulting context-aware PoC attains a superior overall performance. Our work paves the way for a more reliable, efficient, and performance-aware deployment of context compression for LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19733
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction
Zhao, Runsong
Liu, Shilei
Tang, Jiwei
Liu, Langming
Chen, Haibin
Zhang, Weidong
Yuan, Yujin
Xiao, Tong
Zhu, Jingbo
Su, Wenbo
Zheng, Bo
Computation and Language
While context compression can mitigate the growing inference costs of Large Language Models (LLMs) by shortening contexts, existing methods that specify a target compression ratio or length suffer from unpredictable performance degradation, hindering their reliable deployment. We introduce a paradigm shift to Performance-oriented Context Compression (PoC), where developers specify an acceptable performance floor instead of a compression ratio. PoC employs a lightweight performance predictor to automatically find the most aggressive compression ratio that satisfies this constraint before steering an off-the-shelf compressor. We design and compare two predictor variants: a simple context-agnostic predictor and a more sophisticated context-aware one that considers the input's inherent compressibility. On both question-answering and summarization benchmarks, the context-aware predictor consistently achieves lower performance prediction error than the context-agnostic predictor, while the resulting context-aware PoC attains a superior overall performance. Our work paves the way for a more reliable, efficient, and performance-aware deployment of context compression for LLMs.
title PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction
topic Computation and Language
url https://arxiv.org/abs/2603.19733