Concept than Document: Context Compression via AMR-based Conceptual Entropy

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Kaize, Sun, Xueyao, Tao, Xiaohui, Li, Lin, Lin, Qika, Xu, Guandong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911283509460992
author Shi, Kaize
Sun, Xueyao
Tao, Xiaohui
Li, Lin
Lin, Qika
Xu, Guandong
author_facet Shi, Kaize
Sun, Xueyao
Tao, Xiaohui
Li, Lin
Lin, Qika
Xu, Guandong
contents Large Language Models (LLMs) face information overload when handling long contexts, particularly in Retrieval-Augmented Generation (RAG) where extensive supporting documents often introduce redundant content. This issue not only weakens reasoning accuracy but also increases computational overhead. We propose an unsupervised context compression framework that exploits Abstract Meaning Representation (AMR) graphs to preserve semantically essential information while filtering out irrelevant text. By quantifying node-level entropy within AMR graphs, our method estimates the conceptual importance of each node, enabling the retention of core semantics. Specifically, we construct AMR graphs from raw contexts, compute the conceptual entropy of each node, and screen significant informative nodes to form a condensed and semantically focused context than raw documents. Experiments on the PopQA and EntityQuestions datasets show that our method outperforms vanilla and other baselines, achieving higher accuracy while substantially reducing context length. To the best of our knowledge, this is the first work introducing AMR-based conceptual entropy for context compression, demonstrating the potential of stable linguistic features in context engineering.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concept than Document: Context Compression via AMR-based Conceptual Entropy
Shi, Kaize
Sun, Xueyao
Tao, Xiaohui
Li, Lin
Lin, Qika
Xu, Guandong
Computation and Language
Large Language Models (LLMs) face information overload when handling long contexts, particularly in Retrieval-Augmented Generation (RAG) where extensive supporting documents often introduce redundant content. This issue not only weakens reasoning accuracy but also increases computational overhead. We propose an unsupervised context compression framework that exploits Abstract Meaning Representation (AMR) graphs to preserve semantically essential information while filtering out irrelevant text. By quantifying node-level entropy within AMR graphs, our method estimates the conceptual importance of each node, enabling the retention of core semantics. Specifically, we construct AMR graphs from raw contexts, compute the conceptual entropy of each node, and screen significant informative nodes to form a condensed and semantically focused context than raw documents. Experiments on the PopQA and EntityQuestions datasets show that our method outperforms vanilla and other baselines, achieving higher accuracy while substantially reducing context length. To the best of our knowledge, this is the first work introducing AMR-based conceptual entropy for context compression, demonstrating the potential of stable linguistic features in context engineering.
title Concept than Document: Context Compression via AMR-based Conceptual Entropy
topic Computation and Language
url https://arxiv.org/abs/2511.18832