Attn-GS: Attention-Guided Context Compression for Efficient Personalized LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Shenglai, Zheng, Tianqi, Tian, Chuan, Everaert, Dante, Wang, Yau-Shian, Huang, Yupin, Morais, Michael J., Patki, Rohit, Tian, Jinjin, Dai, Xinnan, Guo, Kai, Cheng, Monica Xiao, Liu, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset
by: Everaert, Dante, et al.
Published: (2024)
by: Everaert, Dante, et al.
Published: (2024)
Retrieval Augmented Spelling Correction for E-Commerce Applications
by: Guo, Xuan, et al.
Published: (2024)
by: Guo, Xuan, et al.
Published: (2024)
AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation
by: Luo, Lvzhou, et al.
Published: (2025)
by: Luo, Lvzhou, et al.
Published: (2025)
When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression
by: Dai, Xinnan, et al.
Published: (2026)
by: Dai, Xinnan, et al.
Published: (2026)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
Towards Context-Robust LLMs: A Gated Representation Fine-tuning Approach
by: Zeng, Shenglai, et al.
Published: (2025)
by: Zeng, Shenglai, et al.
Published: (2025)
GIO: Gradient Information Optimization for Training Dataset Selection
by: Everaert, Dante, et al.
Published: (2023)
by: Everaert, Dante, et al.
Published: (2023)
Towards Knowledge Checking in Retrieval-augmented Generation: A Representation Perspective
by: Zeng, Shenglai, et al.
Published: (2024)
by: Zeng, Shenglai, et al.
Published: (2024)
Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
by: Dai, Xinnan, et al.
Published: (2025)
by: Dai, Xinnan, et al.
Published: (2025)
LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning
by: Ji, Mengmeng, et al.
Published: (2026)
by: Ji, Mengmeng, et al.
Published: (2026)
ReAttn: Improving Attention-based Re-ranking via Attention Re-weighting
by: Tian, Yuxing, et al.
Published: (2026)
by: Tian, Yuxing, et al.
Published: (2026)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
by: Wang, Zijun, et al.
Published: (2024)
by: Wang, Zijun, et al.
Published: (2024)
ProxyAttn: Guided Sparse Attention via Representative Heads
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Beyond Static Retrieval: Opportunities and Pitfalls of Iterative Retrieval in GraphRAG
by: Guo, Kai, et al.
Published: (2025)
by: Guo, Kai, et al.
Published: (2025)
Fix Before Search: Benchmarking Agentic Query Visual Pre-processing in Multimodal Retrieval-augmented Generation
by: Zhang, Jiankun, et al.
Published: (2026)
by: Zhang, Jiankun, et al.
Published: (2026)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
GraphGhost: Tracing Structures Behind Large Language Models
by: Dai, Xinnan, et al.
Published: (2025)
by: Dai, Xinnan, et al.
Published: (2025)
SpecAttn: Speculating Sparse Attention
by: Shah, Harsh
Published: (2025)
by: Shah, Harsh
Published: (2025)
AttnGen: Attention-Guided Saliency Learning for Interpretable Genomic Sequence Classification
by: Nia, Rayhaneh Shabani, et al.
Published: (2026)
by: Nia, Rayhaneh Shabani, et al.
Published: (2026)
Why Retrieval-Augmented Generation Fails: A Graph Perspective
by: Guo, Kai, et al.
Published: (2026)
by: Guo, Kai, et al.
Published: (2026)
$Δ$-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation
by: Hu, Jucheng, et al.
Published: (2025)
by: Hu, Jucheng, et al.
Published: (2025)
AttnMod: Attention-Based New Art Styles
by: Su, Shih-Chieh
Published: (2024)
by: Su, Shih-Chieh
Published: (2024)
Generation-driven Contrastive Self-training for Zero-shot Text Classification with Instruction-following LLM
by: Zhang, Ruohong, et al.
Published: (2023)
by: Zhang, Ruohong, et al.
Published: (2023)
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
by: Zhang, Hongyu, et al.
Published: (2025)
by: Zhang, Hongyu, et al.
Published: (2025)
D-Attn: Decomposed Attention for Large Vision-and-Language Models
by: Kuo, Chia-Wen, et al.
Published: (2025)
by: Kuo, Chia-Wen, et al.
Published: (2025)
Attn-JGNN: Attention Enhanced Join-Graph Neural Networks
by: Zhang, Jixin
Published: (2025)
by: Zhang, Jixin
Published: (2025)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
by: Zhang, Peiyuan, et al.
Published: (2026)
by: Zhang, Peiyuan, et al.
Published: (2026)
Exploring Graph Learning Tasks with Pure LLMs: A Comprehensive Benchmark and Investigation
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
by: Song, Dinghong, et al.
Published: (2025)
by: Song, Dinghong, et al.
Published: (2025)
Macht in de metropool
by: Everaert, Janna
Published: (2023)
by: Everaert, Janna
Published: (2023)
In-context KV-Cache Eviction for LLMs via Attention-Gate
by: Zeng, Zihao, et al.
Published: (2024)
by: Zeng, Zihao, et al.
Published: (2024)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
by: Liu, Xuewen, et al.
Published: (2025)
by: Liu, Xuewen, et al.
Published: (2025)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
AttnDiff: Attention-based Differential Fingerprinting for Large Language Models
by: Zhang, Haobo, et al.
Published: (2026)
by: Zhang, Haobo, et al.
Published: (2026)
Pragmatic Study of Speech Act in Beckett’s Waiting for Godot Act - I
by: Sachin L. Patki
Published: (2017)
by: Sachin L. Patki
Published: (2017)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
Evaluation of an Extendable Context-Aware "Learning Java" App with Personalized User Profiling
by: Yau, Jane Yin-Kim, et al.
Published: (2018)
by: Yau, Jane Yin-Kim, et al.
Published: (2018)
TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
by: Zhang, Hongyu, et al.
Published: (2026)
by: Zhang, Hongyu, et al.
Published: (2026)
GradAttn: Replacing Fixed Residual Connections with Task-Modulated Attention Pathways
by: Ghoshal, Soudeep, et al.
Published: (2026)
by: Ghoshal, Soudeep, et al.
Published: (2026)
Similar Items
-
AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset
by: Everaert, Dante, et al.
Published: (2024) -
Retrieval Augmented Spelling Correction for E-Commerce Applications
by: Guo, Xuan, et al.
Published: (2024) -
AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation
by: Luo, Lvzhou, et al.
Published: (2025) -
When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression
by: Dai, Xinnan, et al.
Published: (2026) -
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)