Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Lv, Kangtao, Tang, Jiwei, Liu, Langming, Chen, Haibin, Zhang, Weidong, Liu, Shilei, Wang, Yongwei, Yuan, Yujin, Su, Wenbo, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
by: Lv, Kangtao, et al.
Published: (2025)
by: Lv, Kangtao, et al.
Published: (2025)
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
by: Liu, Langming, et al.
Published: (2026)
by: Liu, Langming, et al.
Published: (2026)
PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
by: Chen, Haibin, et al.
Published: (2025)
by: Chen, Haibin, et al.
Published: (2025)
Read As Human: Compressing Context via Parallelizable Close Reading and Skimming
by: Tang, Jiwei, et al.
Published: (2026)
by: Tang, Jiwei, et al.
Published: (2026)
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
ECKGBench: Benchmarking Large Language Models in E-commerce Leveraging Knowledge Graph
by: Liu, Langming, et al.
Published: (2025)
by: Liu, Langming, et al.
Published: (2025)
COMI: Coarse-to-fine Context Compression via Marginal Information Gain
by: Tang, Jiwei, et al.
Published: (2026)
by: Tang, Jiwei, et al.
Published: (2026)
Expert Divergence Learning for MoE-based Language Models
by: Li, Jiaang, et al.
Published: (2026)
by: Li, Jiaang, et al.
Published: (2026)
Unlocking Scaling Law in Industrial Recommendation Systems with a Three-step Paradigm based Large User Model
by: Yan, Bencheng, et al.
Published: (2025)
by: Yan, Bencheng, et al.
Published: (2025)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
UQABench: Evaluating User Embedding for Prompting LLMs in Personalized Question Answering
by: Liu, Langming, et al.
Published: (2025)
by: Liu, Langming, et al.
Published: (2025)
Hyper Adversarial Tuning for Boosting Adversarial Robustness of Pretrained Large Vision Models
by: Lv, Kangtao, et al.
Published: (2024)
by: Lv, Kangtao, et al.
Published: (2024)
HyperDet: Generalizable Detection of Synthesized Images by Generating and Merging A Mixture of Hyper LoRAs
by: Cao, Huangsen, et al.
Published: (2024)
by: Cao, Huangsen, et al.
Published: (2024)
Multi-task Offline Reinforcement Learning for Online Advertising in Recommender Systems
by: Liu, Langming, et al.
Published: (2025)
by: Liu, Langming, et al.
Published: (2025)
Shifting AI Efficiency From Model-Centric to Data-Centric Compression
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
Analysis of regularized federated learning
by: Liu, Langming, et al.
Published: (2024)
by: Liu, Langming, et al.
Published: (2024)
Position IDs Matter: An Enhanced Position Layout for Efficient Context Compression in Large Language Models
by: Zhao, Runsong, et al.
Published: (2024)
by: Zhao, Runsong, et al.
Published: (2024)
Naive Bayes-based Context Extension for Large Language Models
by: Su, Jianlin, et al.
Published: (2024)
by: Su, Jianlin, et al.
Published: (2024)
Visual Text Compression as Measure Transport
by: Tang, Lv, et al.
Published: (2026)
by: Tang, Lv, et al.
Published: (2026)
ProgCo: Program Helps Self-Correction of Large Language Models
by: Song, Xiaoshuai, et al.
Published: (2025)
by: Song, Xiaoshuai, et al.
Published: (2025)
Artificial Data, Real Insights: Evaluating Opportunities and Risks of Expanding the Data Ecosystem with Synthetic Data
by: Timpone, Richard, et al.
Published: (2024)
by: Timpone, Richard, et al.
Published: (2024)
Data Augmentation in Human-Centric Vision
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
Data-Centric AI in the Age of Large Language Models
by: Xu, Xinyi, et al.
Published: (2024)
by: Xu, Xinyi, et al.
Published: (2024)
Distributed Estimation and Inference for Semi-parametric Binary Response Models
by: Chen, Xi, et al.
Published: (2022)
by: Chen, Xi, et al.
Published: (2022)
GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
Extending Context Window of Large Language Models from a Distributional Perspective
by: Wu, Yingsheng, et al.
Published: (2024)
by: Wu, Yingsheng, et al.
Published: (2024)
PVContext: Hybrid Context Model for Point Cloud Compression
by: Zhang, Guoqing, et al.
Published: (2024)
by: Zhang, Guoqing, et al.
Published: (2024)
Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios
by: Tang, Jiwei, et al.
Published: (2024)
by: Tang, Jiwei, et al.
Published: (2024)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
by: Wu, Yanan, et al.
Published: (2024)
by: Wu, Yanan, et al.
Published: (2024)
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
An Improved Quantum Private Set Intersection Protocol Based on Hadamard Gates
by: Liu, Wenjie, et al.
Published: (2023)
by: Liu, Wenjie, et al.
Published: (2023)
Large Language Models to Diffusion Finetuning
by: Cetin, Edoardo, et al.
Published: (2025)
by: Cetin, Edoardo, et al.
Published: (2025)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven
by: Tang, Jiwei, et al.
Published: (2026)
by: Tang, Jiwei, et al.
Published: (2026)
Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
by: Liu, Hexin, et al.
Published: (2025)
by: Liu, Hexin, et al.
Published: (2025)
Advancements in 3D Lane Detection Using LiDAR Point Clouds: From Data Collection to Model Development
by: Zhao, Runkai, et al.
Published: (2023)
by: Zhao, Runkai, et al.
Published: (2023)
MateICL: Mitigating Attention Dispersion in Large-Scale In-Context Learning
by: Ahmed, Murtadha, et al.
Published: (2025)
by: Ahmed, Murtadha, et al.
Published: (2025)
DataMaster: Data-Centric Autonomous AI Research
by: Du, Yaxin, et al.
Published: (2026)
by: Du, Yaxin, et al.
Published: (2026)
A Data-Driven Modeling and Motion Control of Heavy-Load Hydraulic Manipulators via Reversible Transformation
by: Ma, Dexian, et al.
Published: (2024)
by: Ma, Dexian, et al.
Published: (2024)
Similar Items
-
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
by: Lv, Kangtao, et al.
Published: (2025) -
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
by: Liu, Langming, et al.
Published: (2026) -
PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction
by: Zhao, Runsong, et al.
Published: (2026) -
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
by: Chen, Haibin, et al.
Published: (2025) -
Read As Human: Compressing Context via Parallelizable Close Reading and Skimming
by: Tang, Jiwei, et al.
Published: (2026)