Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
Fuente:
arXiv
Guardado en:
| Autores principales: | Figliolia, Tomas, Alonso, Nicholas, Iyer, Rishi, Anthony, Quentin, Millidge, Beren |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
por: Alonso, Nick, et al.
Publicado: (2024)
por: Alonso, Nick, et al.
Publicado: (2024)
Zyda-2: a 5 Trillion Token High-Quality Dataset
por: Tokpanov, Yury, et al.
Publicado: (2024)
por: Tokpanov, Yury, et al.
Publicado: (2024)
BlackMamba: Mixture of Experts for State-Space Models
por: Anthony, Quentin, et al.
Publicado: (2024)
por: Anthony, Quentin, et al.
Publicado: (2024)
Online Vector Quantized Attention
por: Alonso, Nick, et al.
Publicado: (2026)
por: Alonso, Nick, et al.
Publicado: (2026)
Hybrid Associative Memories
por: Lufkin, Leon, et al.
Publicado: (2026)
por: Lufkin, Leon, et al.
Publicado: (2026)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
por: Shyam, Vasudev, et al.
Publicado: (2024)
por: Shyam, Vasudev, et al.
Publicado: (2024)
ZAYA1-8B Technical Report
por: Washbourne, Robert, et al.
Publicado: (2026)
por: Washbourne, Robert, et al.
Publicado: (2026)
Zyda: A 1.3T Dataset for Open Language Modeling
por: Tokpanov, Yury, et al.
Publicado: (2024)
por: Tokpanov, Yury, et al.
Publicado: (2024)
Zamba: A Compact 7B SSM Hybrid Model
por: Glorioso, Paolo, et al.
Publicado: (2024)
por: Glorioso, Paolo, et al.
Publicado: (2024)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
por: Saxena, Utkarsh, et al.
Publicado: (2024)
por: Saxena, Utkarsh, et al.
Publicado: (2024)
The Zamba2 Suite: Technical Report
por: Glorioso, Paolo, et al.
Publicado: (2024)
por: Glorioso, Paolo, et al.
Publicado: (2024)
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
por: Anthony, Quentin, et al.
Publicado: (2025)
por: Anthony, Quentin, et al.
Publicado: (2025)
LatentLLM: Attention-Aware Joint Tensor Compression
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
por: Chen, Lizhe, et al.
Publicado: (2025)
por: Chen, Lizhe, et al.
Publicado: (2025)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
por: Yun, Jungmin, et al.
Publicado: (2024)
por: Yun, Jungmin, et al.
Publicado: (2024)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
por: Devoto, Alessio, et al.
Publicado: (2025)
por: Devoto, Alessio, et al.
Publicado: (2025)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
por: Zhang, Yong, et al.
Publicado: (2025)
por: Zhang, Yong, et al.
Publicado: (2025)
Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models
por: Xu, Zihao, et al.
Publicado: (2026)
por: Xu, Zihao, et al.
Publicado: (2026)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
por: Wang, Weilan, et al.
Publicado: (2025)
por: Wang, Weilan, et al.
Publicado: (2025)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
por: Hong, Junyuan, et al.
Publicado: (2024)
por: Hong, Junyuan, et al.
Publicado: (2024)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
por: Deniz, Omer Faruk, et al.
Publicado: (2026)
por: Deniz, Omer Faruk, et al.
Publicado: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
por: Yang, Dongquan, et al.
Publicado: (2025)
por: Yang, Dongquan, et al.
Publicado: (2025)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
por: Li, Mengjie, et al.
Publicado: (2025)
por: Li, Mengjie, et al.
Publicado: (2025)
Compressible Softmax-Attended Language under Incompressible Attention
por: Lee, Wonsuk
Publicado: (2026)
por: Lee, Wonsuk
Publicado: (2026)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
por: Huang, Ruizhe, et al.
Publicado: (2025)
por: Huang, Ruizhe, et al.
Publicado: (2025)
Latent Multi-Head Attention for Small Language Models
por: Mehta, Sushant, et al.
Publicado: (2025)
por: Mehta, Sushant, et al.
Publicado: (2025)
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
por: Liu, Zeyu, et al.
Publicado: (2025)
por: Liu, Zeyu, et al.
Publicado: (2025)
Projected Compression: Trainable Projection for Efficient Transformer Compression
por: Stefaniak, Maciej, et al.
Publicado: (2025)
por: Stefaniak, Maciej, et al.
Publicado: (2025)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
por: Knupp, Jonas, et al.
Publicado: (2026)
por: Knupp, Jonas, et al.
Publicado: (2026)
EFPC: Towards Efficient and Flexible Prompt Compression
por: Cao, Yun-Hao, et al.
Publicado: (2025)
por: Cao, Yun-Hao, et al.
Publicado: (2025)
LoRA-Squeeze: Simple and Effective Post-Tuning and In-Tuning Compression of LoRA Modules
por: Vulić, Ivan, et al.
Publicado: (2026)
por: Vulić, Ivan, et al.
Publicado: (2026)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
por: S, Santhosh G, et al.
Publicado: (2025)
por: S, Santhosh G, et al.
Publicado: (2025)
ConvD: Attention Enhanced Dynamic Convolutional Embeddings for Knowledge Graph Completion
por: Guo, Wenbin, et al.
Publicado: (2023)
por: Guo, Wenbin, et al.
Publicado: (2023)
RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step
por: Luo, Xiaocheng, et al.
Publicado: (2026)
por: Luo, Xiaocheng, et al.
Publicado: (2026)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
por: Yankun, Hong, et al.
Publicado: (2025)
por: Yankun, Hong, et al.
Publicado: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
por: Van Nguyen, Chien, et al.
Publicado: (2024)
por: Van Nguyen, Chien, et al.
Publicado: (2024)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
por: Hu, Jinwu, et al.
Publicado: (2025)
por: Hu, Jinwu, et al.
Publicado: (2025)
Recurrent Context Compression: Efficiently Expanding the Context Window of LLM
por: Huang, Chensen, et al.
Publicado: (2024)
por: Huang, Chensen, et al.
Publicado: (2024)
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
por: Yang, Ge, et al.
Publicado: (2024)
por: Yang, Ge, et al.
Publicado: (2024)
Efficient Streaming Language Models with Attention Sinks
por: Xiao, Guangxuan, et al.
Publicado: (2023)
por: Xiao, Guangxuan, et al.
Publicado: (2023)
Ejemplares similares
-
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
por: Alonso, Nick, et al.
Publicado: (2024) -
Zyda-2: a 5 Trillion Token High-Quality Dataset
por: Tokpanov, Yury, et al.
Publicado: (2024) -
BlackMamba: Mixture of Experts for State-Space Models
por: Anthony, Quentin, et al.
Publicado: (2024) -
Online Vector Quantized Attention
por: Alonso, Nick, et al.
Publicado: (2026) -
Hybrid Associative Memories
por: Lufkin, Leon, et al.
Publicado: (2026)