SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yankun, Hong, Xing, Li, Hui-Ling, Zhen, Xianzhi, Yu, Wulong, Liu, Mingxuan, Yuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
von: Johnson, Warren
Veröffentlicht: (2026)
von: Johnson, Warren
Veröffentlicht: (2026)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
von: Johnson, Warren
Veröffentlicht: (2026)
von: Johnson, Warren
Veröffentlicht: (2026)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
Pay Attention to What You Need
von: Gao, Yifei, et al.
Veröffentlicht: (2023)
von: Gao, Yifei, et al.
Veröffentlicht: (2023)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
von: Otal, Hakan T., et al.
Veröffentlicht: (2024)
von: Otal, Hakan T., et al.
Veröffentlicht: (2024)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
von: Chen, Jiaju, et al.
Veröffentlicht: (2025)
von: Chen, Jiaju, et al.
Veröffentlicht: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
von: Dahlem, Dominik, et al.
Veröffentlicht: (2026)
von: Dahlem, Dominik, et al.
Veröffentlicht: (2026)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
von: Yao, Huaiyuan, et al.
Veröffentlicht: (2026)
von: Yao, Huaiyuan, et al.
Veröffentlicht: (2026)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
von: Johnson, Warren, et al.
Veröffentlicht: (2026)
von: Johnson, Warren, et al.
Veröffentlicht: (2026)
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
von: Harris, Joshua, et al.
Veröffentlicht: (2025)
von: Harris, Joshua, et al.
Veröffentlicht: (2025)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
von: Liu, Linyu, et al.
Veröffentlicht: (2024)
von: Liu, Linyu, et al.
Veröffentlicht: (2024)
Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations
von: Bhattacharya, Shamik, et al.
Veröffentlicht: (2026)
von: Bhattacharya, Shamik, et al.
Veröffentlicht: (2026)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2024)
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
von: Andrylie, Lyzander Marciano, et al.
Veröffentlicht: (2025)
von: Andrylie, Lyzander Marciano, et al.
Veröffentlicht: (2025)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
von: Land, Sander, et al.
Veröffentlicht: (2024)
von: Land, Sander, et al.
Veröffentlicht: (2024)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
von: Ikoma, Hayato, et al.
Veröffentlicht: (2025)
von: Ikoma, Hayato, et al.
Veröffentlicht: (2025)
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2023)
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2023)
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
von: Li, Yongchang, et al.
Veröffentlicht: (2024)
von: Li, Yongchang, et al.
Veröffentlicht: (2024)
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
von: Alkan, Atilla Kaan, et al.
Veröffentlicht: (2025)
von: Alkan, Atilla Kaan, et al.
Veröffentlicht: (2025)
Which Pieces Does Unigram Tokenization Really Need?
von: Land, Sander, et al.
Veröffentlicht: (2025)
von: Land, Sander, et al.
Veröffentlicht: (2025)
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2026)
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2026)
Duluth at SemEval-2025 Task 7: TF-IDF with Optimized Vector Dimensions for Multilingual Fact-Checked Claim Retrieval
von: Syed, Shujauddin, et al.
Veröffentlicht: (2025)
von: Syed, Shujauddin, et al.
Veröffentlicht: (2025)
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
IteRABRe: Iterative Recovery-Aided Block Reduction
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2025)
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
von: AlquBoj, H. V., et al.
Veröffentlicht: (2025)
von: AlquBoj, H. V., et al.
Veröffentlicht: (2025)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
von: Seo, Yeongbin, et al.
Veröffentlicht: (2024)
von: Seo, Yeongbin, et al.
Veröffentlicht: (2024)
Surfing the modeling of PoS taggers in low-resource scenarios
von: Ferro, Manuel Vilares, et al.
Veröffentlicht: (2024)
von: Ferro, Manuel Vilares, et al.
Veröffentlicht: (2024)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
Curriculum Recommendations Using Transformer Base Model with InfoNCE Loss And Language Switching Method
von: Xu, Xiaonan, et al.
Veröffentlicht: (2024)
von: Xu, Xiaonan, et al.
Veröffentlicht: (2024)
The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
von: Johnson, Warren
Veröffentlicht: (2026) -
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025) -
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
von: Johnson, Warren
Veröffentlicht: (2026) -
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
von: Basu, Abhinaba
Veröffentlicht: (2026) -
Pay Attention to What You Need
von: Gao, Yifei, et al.
Veröffentlicht: (2023)