Read As Human: Compressing Context via Parallelizable Close Reading and Skimming

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Jiwei, Liu, Shilei, Zhang, Zhicheng, Lv, Qingsong, Zhao, Runsong, Lu, Tingwei, Liu, Langming, Chen, Haibin, Yuan, Yujin, Zheng, Hai-Tao, Su, Wenbo, Zheng, Bo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910034601967616
author Tang, Jiwei
Liu, Shilei
Zhang, Zhicheng
Lv, Qingsong
Zhao, Runsong
Lu, Tingwei
Liu, Langming
Chen, Haibin
Yuan, Yujin
Zheng, Hai-Tao
Su, Wenbo
Zheng, Bo
author_facet Tang, Jiwei
Liu, Shilei
Zhang, Zhicheng
Lv, Qingsong
Zhao, Runsong
Lu, Tingwei
Liu, Langming
Chen, Haibin
Yuan, Yujin
Zheng, Hai-Tao
Su, Wenbo
Zheng, Bo
contents Large Language Models (LLMs) demonstrate exceptional capability across diverse tasks. However, their deployment in long-context scenarios is hindered by two challenges: computational inefficiency and redundant information. We propose RAM (Read As HuMan), a context compression framework that adopts an adaptive hybrid reading strategy, to address these challenges. Inspired by human reading behavior (i.e., close reading important content while skimming less relevant content), RAM partitions the context into segments and encodes them with the input query in parallel. High-relevance segments are fully retained (close reading), while low-relevance ones are query-guided compressed into compact summary vectors (skimming). Both explicit textual segments and implicit summary vectors are concatenated and fed into decoder to achieve both superior performance and natural language format interpretability. To refine the decision boundary between close reading and skimming, we further introduce a contrastive learning objective based on positive and negative query-segment pairs. Experiments demonstrate that RAM outperforms existing baselines on multiple question answering and summarization benchmarks across two backbones, while delivering up to a 12x end-to-end speedup on long inputs (average length 16K; maximum length 32K).
format Preprint
id arxiv_https___arxiv_org_abs_2602_01840
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Read As Human: Compressing Context via Parallelizable Close Reading and Skimming
Tang, Jiwei
Liu, Shilei
Zhang, Zhicheng
Lv, Qingsong
Zhao, Runsong
Lu, Tingwei
Liu, Langming
Chen, Haibin
Yuan, Yujin
Zheng, Hai-Tao
Su, Wenbo
Zheng, Bo
Computation and Language
Large Language Models (LLMs) demonstrate exceptional capability across diverse tasks. However, their deployment in long-context scenarios is hindered by two challenges: computational inefficiency and redundant information. We propose RAM (Read As HuMan), a context compression framework that adopts an adaptive hybrid reading strategy, to address these challenges. Inspired by human reading behavior (i.e., close reading important content while skimming less relevant content), RAM partitions the context into segments and encodes them with the input query in parallel. High-relevance segments are fully retained (close reading), while low-relevance ones are query-guided compressed into compact summary vectors (skimming). Both explicit textual segments and implicit summary vectors are concatenated and fed into decoder to achieve both superior performance and natural language format interpretability. To refine the decision boundary between close reading and skimming, we further introduce a contrastive learning objective based on positive and negative query-segment pairs. Experiments demonstrate that RAM outperforms existing baselines on multiple question answering and summarization benchmarks across two backbones, while delivering up to a 12x end-to-end speedup on long inputs (average length 16K; maximum length 32K).
title Read As Human: Compressing Context via Parallelizable Close Reading and Skimming
topic Computation and Language
url https://arxiv.org/abs/2602.01840