From Similarity to Structure: Training-free LLM Context Compression with Hybrid Graph Priors

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Yitian, Zhang, Chaoning, Zhang, Jiaquan, Huang, Zhenzhen, Guo, Jinyu, Bae, Sung-Ho, Lee, Lik-Hang, Qin, Caiyan, Yang, Yang
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918467971579904
author Zhou, Yitian
Zhang, Chaoning
Zhang, Jiaquan
Huang, Zhenzhen
Guo, Jinyu
Bae, Sung-Ho
Lee, Lik-Hang
Qin, Caiyan
Yang, Yang
author_facet Zhou, Yitian
Zhang, Chaoning
Zhang, Jiaquan
Huang, Zhenzhen
Guo, Jinyu
Bae, Sung-Ho
Lee, Lik-Hang
Qin, Caiyan
Yang, Yang
contents Long-context large language models remain computationally expensive to run and often fail to reliably process very long inputs, which makes context compression an important component of many systems. Existing compression approaches typically rely on trained compressors, dense retrieval-style selection, or heuristic trimming, and they often struggle to jointly preserve task relevance, topic coverage, and cross-sentence coherence under a strict token budget. To address this, we propose a training-free and model-agnostic compression framework that selects a compact set of sentences guided by structural graph priors. Our method constructs a sparse hybrid sentence graph that combines mutual k-NN semantic edges with short-range sequential edges, extracts a topic skeleton via clustering, and ranks sentences using an interpretable score that integrates task relevance, cluster representativeness, bridge centrality, and a cycle coverage cue. A budgeted greedy selection with redundancy suppression then produces a readable compressed context in original order. Experimental results on four datasets show that our approach is competitive with strong extractive and abstractive baselines, demonstrating larger gains on long-document benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23277
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Similarity to Structure: Training-free LLM Context Compression with Hybrid Graph Priors
Zhou, Yitian
Zhang, Chaoning
Zhang, Jiaquan
Huang, Zhenzhen
Guo, Jinyu
Bae, Sung-Ho
Lee, Lik-Hang
Qin, Caiyan
Yang, Yang
Computation and Language
Artificial Intelligence
Long-context large language models remain computationally expensive to run and often fail to reliably process very long inputs, which makes context compression an important component of many systems. Existing compression approaches typically rely on trained compressors, dense retrieval-style selection, or heuristic trimming, and they often struggle to jointly preserve task relevance, topic coverage, and cross-sentence coherence under a strict token budget. To address this, we propose a training-free and model-agnostic compression framework that selects a compact set of sentences guided by structural graph priors. Our method constructs a sparse hybrid sentence graph that combines mutual k-NN semantic edges with short-range sequential edges, extracts a topic skeleton via clustering, and ranks sentences using an interpretable score that integrates task relevance, cluster representativeness, bridge centrality, and a cycle coverage cue. A budgeted greedy selection with redundancy suppression then produces a readable compressed context in original order. Experimental results on four datasets show that our approach is competitive with strong extractive and abstractive baselines, demonstrating larger gains on long-document benchmarks.
title From Similarity to Structure: Training-free LLM Context Compression with Hybrid Graph Priors
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.23277