ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, David H., Zhu, Yuxuan, Amiri, Mohammad Mohammadi, Murugesan, Keerthiram, Pedapati, Tejaswini, Chaudhury, Subhajit, Chen, Pin-Yu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Sparse Gradient Compression for Fine-Tuning Large Language Models
by: Yang, David H., et al.
Published: (2025)
by: Yang, David H., et al.
Published: (2025)
On the Effects of Fine-tuning Language Models for Text-Based Reinforcement Learning
by: Gruppi, Mauricio, et al.
Published: (2024)
by: Gruppi, Mauricio, et al.
Published: (2024)
EXPLORER: Exploration-guided Reasoning for Textual Reinforcement Learning
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Stochastic Approach
by: Fernando, Heshan, et al.
Published: (2022)
by: Fernando, Heshan, et al.
Published: (2022)
From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in Transformers
by: Runwal, Bharat, et al.
Published: (2024)
by: Runwal, Bharat, et al.
Published: (2024)
EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts
by: Chaudhury, Subhajit, et al.
Published: (2025)
by: Chaudhury, Subhajit, et al.
Published: (2025)
Large Language Models can be Strong Self-Detoxifiers
by: Ko, Ching-Yun, et al.
Published: (2024)
by: Ko, Ching-Yun, et al.
Published: (2024)
LongFuncEval: Measuring the effectiveness of long context models for function calling
by: Kate, Kiran, et al.
Published: (2025)
by: Kate, Kiran, et al.
Published: (2025)
Modular Prompt Learning Improves Vision-Language Models
by: Huang, Zhenhan, et al.
Published: (2025)
by: Huang, Zhenhan, et al.
Published: (2025)
Differentiable Prompt Learning for Vision Language Models
by: Huang, Zhenhan, et al.
Published: (2024)
by: Huang, Zhenhan, et al.
Published: (2024)
Intermediate Representations are Strong AI-Generated Image Detectors
by: Huang, Zhenhan, et al.
Published: (2026)
by: Huang, Zhenhan, et al.
Published: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
by: Huang, Zhenhan, et al.
Published: (2024)
by: Huang, Zhenhan, et al.
Published: (2024)
TabSketchFM: Sketch-based Tabular Representation Learning for Data Discovery over Data Lakes
by: Khatiwada, Aamod, et al.
Published: (2024)
by: Khatiwada, Aamod, et al.
Published: (2024)
PEEL the Layers and Find Yourself: Revisiting Inference-time Data Leakage for Residual Neural Networks
by: Arif, Huzaifa, et al.
Published: (2025)
by: Arif, Huzaifa, et al.
Published: (2025)
STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models
by: Basavatia, Shreyas, et al.
Published: (2024)
by: Basavatia, Shreyas, et al.
Published: (2024)
Toward Efficient Influence Function: Dropout as a Compression Tool
by: Zhang, Yuchen, et al.
Published: (2025)
by: Zhang, Yuchen, et al.
Published: (2025)
Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes
by: Alipour, Mohammadsajad, et al.
Published: (2025)
by: Alipour, Mohammadsajad, et al.
Published: (2025)
Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems
by: Asif, Sadia, et al.
Published: (2026)
by: Asif, Sadia, et al.
Published: (2026)
Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
by: Das, Payel, et al.
Published: (2025)
by: Das, Payel, et al.
Published: (2025)
Context Attribution with Multi-Armed Bandit Optimization
by: Pan, Deng, et al.
Published: (2025)
by: Pan, Deng, et al.
Published: (2025)
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
by: Mukherjee, Arpan, et al.
Published: (2024)
by: Mukherjee, Arpan, et al.
Published: (2024)
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024)
by: Nelson, Elliot, et al.
Published: (2024)
Targeted Advertising on Social Networks Using Online Variational Tensor Regression
by: Idé, Tsuyoshi, et al.
Published: (2022)
by: Idé, Tsuyoshi, et al.
Published: (2022)
Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
by: Arif, Huzaifa, et al.
Published: (2025)
by: Arif, Huzaifa, et al.
Published: (2025)
Large Language Model Confidence Estimation via Black-Box Access
by: Pedapati, Tejaswini, et al.
Published: (2024)
by: Pedapati, Tejaswini, et al.
Published: (2024)
Generation Constraint Scaling Can Mitigate Hallucination
by: Kollias, Georgios, et al.
Published: (2024)
by: Kollias, Georgios, et al.
Published: (2024)
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
by: Gourabathina, Abinitha, et al.
Published: (2026)
by: Gourabathina, Abinitha, et al.
Published: (2026)
Towards Aligning Language Models with Textual Feedback
by: Lloret, Saüc Abadal, et al.
Published: (2024)
by: Lloret, Saüc Abadal, et al.
Published: (2024)
TOPJoin: A Context-Aware Multi-Criteria Approach for Joinable Column Search
by: Kokel, Harsha, et al.
Published: (2025)
by: Kokel, Harsha, et al.
Published: (2025)
STAR: Spectral Truncation and Rescale for Model Merging
by: Lee, Yu-Ang, et al.
Published: (2025)
by: Lee, Yu-Ang, et al.
Published: (2025)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
by: Asif, Sadia, et al.
Published: (2026)
by: Asif, Sadia, et al.
Published: (2026)
OFMU: Optimization-Driven Framework for Machine Unlearning
by: Asif, Sadia, et al.
Published: (2025)
by: Asif, Sadia, et al.
Published: (2025)
Disentangled Structural and Featural Representation for Task-Agnostic Graph Valuation
by: Falahati, Ali, et al.
Published: (2024)
by: Falahati, Ali, et al.
Published: (2024)
Towards Reversible Model Merging For Low-rank Weights
by: Alipour, Mohammadsajad, et al.
Published: (2025)
by: Alipour, Mohammadsajad, et al.
Published: (2025)
Power to the Clients: Federated Learning in a Dictatorship Setting
by: Alipour, Mohammadsajad, et al.
Published: (2025)
by: Alipour, Mohammadsajad, et al.
Published: (2025)
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
R3-REC: Reasoning-Driven Recommendation via Retrieval-Augmented LLMs over Multi-Granular Interest Signals
by: Miao, Yuchen, et al.
Published: (2026)
by: Miao, Yuchen, et al.
Published: (2026)
Memory-Efficient Community Detection on Large Graphs Using Weighted Sketches
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
Similar Items
-
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
by: Zhu, Yuxuan, et al.
Published: (2025) -
Sparse Gradient Compression for Fine-Tuning Large Language Models
by: Yang, David H., et al.
Published: (2025) -
On the Effects of Fine-tuning Language Models for Text-Based Reinforcement Learning
by: Gruppi, Mauricio, et al.
Published: (2024) -
EXPLORER: Exploration-guided Reasoning for Textual Reinforcement Learning
by: Basu, Kinjal, et al.
Published: (2024) -
Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Stochastic Approach
by: Fernando, Heshan, et al.
Published: (2022)