Saved in:
| Main Authors: | Mao, Yuzhen, Li, Michael Y., Fox, Emily B. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.20920 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long Context In-Context Compression by Getting to the Gist of Gisting
by: Petrov, Aleksandar, et al.
Published: (2025)
by: Petrov, Aleksandar, et al.
Published: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025)
by: He, Mutian, et al.
Published: (2025)
Neural Garbage Collection: Learning to Forget while Learning to Reason
by: Li, Michael Y., et al.
Published: (2026)
by: Li, Michael Y., et al.
Published: (2026)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
by: Piękos, Piotr, et al.
Published: (2025)
by: Piękos, Piotr, et al.
Published: (2025)
Sparse graphs using exchangeable random measures
by: Caron, François, et al.
Published: (2014)
by: Caron, François, et al.
Published: (2014)
Automated Statistical Model Discovery with Language Models
by: Li, Michael Y., et al.
Published: (2024)
by: Li, Michael Y., et al.
Published: (2024)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
SALS: Sparse Attention in Latent Space for KV cache Compression
by: Mu, Junlin, et al.
Published: (2025)
by: Mu, Junlin, et al.
Published: (2025)
IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
by: Mao, Yuzhen, et al.
Published: (2024)
by: Mao, Yuzhen, et al.
Published: (2024)
Tuning-Free Structured Sparse PCA via Deep Unfolding Networks
by: Chen, Long, et al.
Published: (2025)
by: Chen, Long, et al.
Published: (2025)
Representation Unlearning: Forgetting through Information Compression
by: Almudévar, Antonio, et al.
Published: (2026)
by: Almudévar, Antonio, et al.
Published: (2026)
Gated Graph Attention Networks with Learnable Temperature
by: Ma, Zhongtian, et al.
Published: (2026)
by: Ma, Zhongtian, et al.
Published: (2026)
Learnability in Online Kernel Selection with Memory Constraint via Data-dependent Regret Analysis
by: Li, Junfan, et al.
Published: (2024)
by: Li, Junfan, et al.
Published: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Domain Adaptation and Multi-view Attention for Learnable Landmark Tracking with Sparse Data
by: Chase Jr, Timothy, et al.
Published: (2025)
by: Chase Jr, Timothy, et al.
Published: (2025)
CriticAL: Critic Automation with Language Models
by: Li, Michael Y., et al.
Published: (2024)
by: Li, Michael Y., et al.
Published: (2024)
A Two-Phase Recall-and-Select Framework for Fast Model Selection
by: Cui, Jianwei, et al.
Published: (2024)
by: Cui, Jianwei, et al.
Published: (2024)
Reflecting Topology Consistency and Abnormality via Learnable Attentions for Airway Labeling
by: Li, Chenyu, et al.
Published: (2024)
by: Li, Chenyu, et al.
Published: (2024)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
by: Shang, Bingqi, et al.
Published: (2025)
by: Shang, Bingqi, et al.
Published: (2025)
Auto-FlexSwitch: Efficient Dynamic Model Merging via Learnable Task Vector Compression
by: Gao, Junqi, et al.
Published: (2026)
by: Gao, Junqi, et al.
Published: (2026)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
Learnable Sparse Customization in Heterogeneous Edge Computing
by: Xue, Jingjing, et al.
Published: (2024)
by: Xue, Jingjing, et al.
Published: (2024)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
by: Yang, Xu, et al.
Published: (2026)
by: Yang, Xu, et al.
Published: (2026)
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
by: Mao, Yuzhen, et al.
Published: (2026)
by: Mao, Yuzhen, et al.
Published: (2026)
Robust Learnability of Sample-Compressible Distributions under Noisy or Adversarial Perturbations
by: Boushehrian, Arefe, et al.
Published: (2025)
by: Boushehrian, Arefe, et al.
Published: (2025)
BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis
by: Ni, Zelin, et al.
Published: (2023)
by: Ni, Zelin, et al.
Published: (2023)
OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
by: Jiang, Liangze, et al.
Published: (2024)
by: Jiang, Liangze, et al.
Published: (2024)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention Dynamics
by: Koubbi, Hugo, et al.
Published: (2024)
by: Koubbi, Hugo, et al.
Published: (2024)
Solving Sparse \& High-Dimensional-Output Regression via Compression
by: Li, Renyuan, et al.
Published: (2024)
by: Li, Renyuan, et al.
Published: (2024)
Performative Risk Control: Calibrating Models for Reliable Deployment under Performativity
by: Li, Victor, et al.
Published: (2025)
by: Li, Victor, et al.
Published: (2025)
Learning to Forget: Bayesian Time Series Forecasting using Recurrent Sparse Spectrum Signature Gaussian Processes
by: Tóth, Csaba, et al.
Published: (2024)
by: Tóth, Csaba, et al.
Published: (2024)
Learning to Forget Attention: Memory Consolidation for Adaptive Compute Reduction
by: Shihab, Ibne Farabi, et al.
Published: (2026)
by: Shihab, Ibne Farabi, et al.
Published: (2026)
Vectorized Attention with Learnable Encoding for Quantum Transformer
by: Guo, Ziqing, et al.
Published: (2025)
by: Guo, Ziqing, et al.
Published: (2025)
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024)
by: Huang, Yuzhen, et al.
Published: (2024)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
by: Gao, Yizhao, et al.
Published: (2025)
by: Gao, Yizhao, et al.
Published: (2025)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
by: Khoriaty, Matthew, et al.
Published: (2025)
by: Khoriaty, Matthew, et al.
Published: (2025)
Experimental Analysis of Large-scale Learnable Vector Storage Compression
by: Zhang, Hailin, et al.
Published: (2023)
by: Zhang, Hailin, et al.
Published: (2023)
Forget the Data and Fine-Tuning! Just Fold the Network to Compress
by: Wang, Dong, et al.
Published: (2025)
by: Wang, Dong, et al.
Published: (2025)
Positional Attention: Expressivity and Learnability of Algorithmic Computation
by: de Luca, Artur Back, et al.
Published: (2024)
by: de Luca, Artur Back, et al.
Published: (2024)
Similar Items
-
Long Context In-Context Compression by Getting to the Gist of Gisting
by: Petrov, Aleksandar, et al.
Published: (2025) -
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025) -
Neural Garbage Collection: Learning to Forget while Learning to Reason
by: Li, Michael Y., et al.
Published: (2026) -
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
by: Piękos, Piotr, et al.
Published: (2025) -
Sparse graphs using exchangeable random measures
by: Caron, François, et al.
Published: (2014)