Guardado en:
| Autores principales: | Belitsky, Max, Kopiczko, Dawid J., Dorkenwald, Michael, Mirza, M. Jehanzeb, Glass, James R., Snoek, Cees G. M., Asano, Yuki M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2507.08799 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
por: Laitenberger, Filipe, et al.
Publicado: (2025)
por: Laitenberger, Filipe, et al.
Publicado: (2025)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
por: Kopiczko, Dawid J., et al.
Publicado: (2024)
por: Kopiczko, Dawid J., et al.
Publicado: (2024)
VeRA: Vector-based Random Matrix Adaptation
por: Kopiczko, Dawid J., et al.
Publicado: (2023)
por: Kopiczko, Dawid J., et al.
Publicado: (2023)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
por: Dorkenwald, Michael, et al.
Publicado: (2024)
por: Dorkenwald, Michael, et al.
Publicado: (2024)
Lost in Time: A New Temporal Benchmark for VideoLLMs
por: Cores, Daniel, et al.
Publicado: (2024)
por: Cores, Daniel, et al.
Publicado: (2024)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
por: Kopiczko, Dawid J., et al.
Publicado: (2026)
por: Kopiczko, Dawid J., et al.
Publicado: (2026)
Elastic ViTs from Pretrained Models without Retraining
por: Simoncini, Walter, et al.
Publicado: (2025)
por: Simoncini, Walter, et al.
Publicado: (2025)
SIGMA: Sinkhorn-Guided Masked Video Modeling
por: Salehi, Mohammadreza, et al.
Publicado: (2024)
por: Salehi, Mohammadreza, et al.
Publicado: (2024)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
por: Rastegar, Sarah, et al.
Publicado: (2024)
por: Rastegar, Sarah, et al.
Publicado: (2024)
NeoBabel: A Multilingual Open Tower for Visual Generation
por: Derakhshani, Mohammad Mahdi, et al.
Publicado: (2025)
por: Derakhshani, Mohammad Mahdi, et al.
Publicado: (2025)
Overflow Prevention Enhances Long-Context Recurrent LLMs
por: Ben-Kish, Assaf, et al.
Publicado: (2025)
por: Ben-Kish, Assaf, et al.
Publicado: (2025)
Beyond Model Adaptation at Test Time: A Survey
por: Xiao, Zehao, et al.
Publicado: (2024)
por: Xiao, Zehao, et al.
Publicado: (2024)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
por: Mehta, Videet, et al.
Publicado: (2026)
por: Mehta, Videet, et al.
Publicado: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
por: Liao, Bingli, et al.
Publicado: (2024)
por: Liao, Bingli, et al.
Publicado: (2024)
Segment Any 3D-Part in a Scene from a Sentence
por: Wu, Hongyu, et al.
Publicado: (2025)
por: Wu, Hongyu, et al.
Publicado: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
por: Nguyen-Tri, Quan, et al.
Publicado: (2025)
por: Nguyen-Tri, Quan, et al.
Publicado: (2025)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
por: Wang, Zihan, et al.
Publicado: (2026)
por: Wang, Zihan, et al.
Publicado: (2026)
IPO: Interpretable Prompt Optimization for Vision-Language Models
por: Du, Yingjun, et al.
Publicado: (2024)
por: Du, Yingjun, et al.
Publicado: (2024)
Lossless KV Cache Compression to 2%
por: Yang, Zhen, et al.
Publicado: (2024)
por: Yang, Zhen, et al.
Publicado: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
por: Liao, Mengqi, et al.
Publicado: (2025)
por: Liao, Mengqi, et al.
Publicado: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
por: Cai, Zefan, et al.
Publicado: (2025)
por: Cai, Zefan, et al.
Publicado: (2025)
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
por: Liu, Andy Zeyi, et al.
Publicado: (2026)
por: Liu, Andy Zeyi, et al.
Publicado: (2026)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
por: Liu, Xin, et al.
Publicado: (2025)
por: Liu, Xin, et al.
Publicado: (2025)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
por: Ghadia, Ravi, et al.
Publicado: (2025)
por: Ghadia, Ravi, et al.
Publicado: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
por: Yang, Dongquan, et al.
Publicado: (2025)
por: Yang, Dongquan, et al.
Publicado: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
por: Cai, Zefan, et al.
Publicado: (2024)
por: Cai, Zefan, et al.
Publicado: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
por: Doughty, Hazel, et al.
Publicado: (2024)
por: Doughty, Hazel, et al.
Publicado: (2024)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
Towards Threshold-Free KV Cache Pruning
por: Ni, Xuanfan, et al.
Publicado: (2025)
por: Ni, Xuanfan, et al.
Publicado: (2025)
KVSculpt: KV Cache Compression as Distillation
por: Jiang, Bo, et al.
Publicado: (2026)
por: Jiang, Bo, et al.
Publicado: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
por: Gupta, Manan, et al.
Publicado: (2026)
por: Gupta, Manan, et al.
Publicado: (2026)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
por: Liu, Huabin, et al.
Publicado: (2025)
por: Liu, Huabin, et al.
Publicado: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
por: Dehghanighobadi, Zahra, et al.
Publicado: (2026)
por: Dehghanighobadi, Zahra, et al.
Publicado: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
por: Hao, Jitai, et al.
Publicado: (2026)
por: Hao, Jitai, et al.
Publicado: (2026)
Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
por: Poudel, Pratik
Publicado: (2025)
por: Poudel, Pratik
Publicado: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024)
por: Feng, Yuan, et al.
Publicado: (2024)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
por: Salehi, Mohammadreza, et al.
Publicado: (2024)
por: Salehi, Mohammadreza, et al.
Publicado: (2024)
KV Cache Offloading for Context-Intensive Tasks
por: Bocharnikov, Andrey, et al.
Publicado: (2026)
por: Bocharnikov, Andrey, et al.
Publicado: (2026)
Ejemplares similares
-
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
por: Laitenberger, Filipe, et al.
Publicado: (2025) -
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
por: Kopiczko, Dawid J., et al.
Publicado: (2024) -
VeRA: Vector-based Random Matrix Adaptation
por: Kopiczko, Dawid J., et al.
Publicado: (2023) -
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
por: Dorkenwald, Michael, et al.
Publicado: (2024) -
Lost in Time: A New Temporal Benchmark for VideoLLMs
por: Cores, Daniel, et al.
Publicado: (2024)