LoMA: Lossless Compressed Memory Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yumeng, Xiao, Zhenyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Llamazip: Leveraging LLaMA for Lossless Text Compression and Training Dataset Detection
von: Dréano, Sören, et al.
Veröffentlicht: (2025)
von: Dréano, Sören, et al.
Veröffentlicht: (2025)
Lossless Token Sequence Compression via Meta-Tokens
von: Harvill, John, et al.
Veröffentlicht: (2025)
von: Harvill, John, et al.
Veröffentlicht: (2025)
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
LoCoCo: Dropping In Convolutions for Long Context Compression
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
Boosting Lossless Speculative Decoding via Feature Sampling and Partial Alignment Distillation
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
von: Mittu, Fazal, et al.
Veröffentlicht: (2024)
von: Mittu, Fazal, et al.
Veröffentlicht: (2024)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database
von: Sun, Hui, et al.
Veröffentlicht: (2025)
von: Sun, Hui, et al.
Veröffentlicht: (2025)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
LoLA: Low-Rank Linear Attention With Sparse Caching
von: McDermott, Luke, et al.
Veröffentlicht: (2025)
von: McDermott, Luke, et al.
Veröffentlicht: (2025)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
Compressed Context Memory For Online Language Model Interaction
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2023)
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2023)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
CompAct: Compressed Activations for Memory-Efficient LLM Training
von: Shamshoum, Yara, et al.
Veröffentlicht: (2024)
von: Shamshoum, Yara, et al.
Veröffentlicht: (2024)
LogLLaMA: Transformer-based log anomaly detection with LLaMA
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2025)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
BiTA: Bi-Directional Tuning for Lossless Acceleration in Large Language Models
von: Lin, Feng, et al.
Veröffentlicht: (2024)
von: Lin, Feng, et al.
Veröffentlicht: (2024)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
von: Gupta, Prakhar, et al.
Veröffentlicht: (2026)
von: Gupta, Prakhar, et al.
Veröffentlicht: (2026)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
Lossless Acceleration of Large Language Model via Adaptive N-gram Parallel Decoding
von: Ou, Jie, et al.
Veröffentlicht: (2024)
von: Ou, Jie, et al.
Veröffentlicht: (2024)
TensorLLM: Tensorising Multi-Head Attention for Enhanced Reasoning and Compression in LLMs
von: Gu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Gu, Yuxuan, et al.
Veröffentlicht: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
Proximity to Losslessly Compressible Parameters
von: Farrugia-Roberts, Matthew
Veröffentlicht: (2023)
von: Farrugia-Roberts, Matthew
Veröffentlicht: (2023)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
von: Chen, Lida, et al.
Veröffentlicht: (2025)
von: Chen, Lida, et al.
Veröffentlicht: (2025)
HyperAdaLoRA: Accelerating LoRA Rank Allocation During Training via Hypernetworks without Sacrificing Performance
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
TeleLoRA: Teleporting Model-Specific Alignment Across LLMs
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Llamazip: Leveraging LLaMA for Lossless Text Compression and Training Dataset Detection
von: Dréano, Sören, et al.
Veröffentlicht: (2025) -
Lossless Token Sequence Compression via Meta-Tokens
von: Harvill, John, et al.
Veröffentlicht: (2025) -
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025) -
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025) -
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)