AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | S, Santhosh G, Prakash, Saurav, Ravindran, Balaraman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics
by: Acharjee, Jashaswimalya, et al.
Published: (2026)
by: Acharjee, Jashaswimalya, et al.
Published: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
by: Yang, Dongquan, et al.
Published: (2025)
by: Yang, Dongquan, et al.
Published: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
by: Butler, Landon, et al.
Published: (2025)
by: Butler, Landon, et al.
Published: (2025)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
by: Waldendorf, Jonas, et al.
Published: (2026)
by: Waldendorf, Jonas, et al.
Published: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024)
by: Deiseroth, Björn, et al.
Published: (2024)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
by: Vendrell, Victor Conchello, et al.
Published: (2026)
by: Vendrell, Victor Conchello, et al.
Published: (2026)
SAGE: A Novelty Gate for Efficient Memory Evolution in Agentic LLMs
by: Wang, Sijia, et al.
Published: (2026)
by: Wang, Sijia, et al.
Published: (2026)
Hallucination Detection in LLMs: Fast and Memory-Efficient Fine-Tuned Models
by: Arteaga, Gabriel Y., et al.
Published: (2024)
by: Arteaga, Gabriel Y., et al.
Published: (2024)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
by: Alizadeh, Keivan, et al.
Published: (2023)
by: Alizadeh, Keivan, et al.
Published: (2023)
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
by: Ravindran, Santhosh Kumar
Published: (2025)
by: Ravindran, Santhosh Kumar
Published: (2025)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
by: Fei, Yu, et al.
Published: (2024)
by: Fei, Yu, et al.
Published: (2024)
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
by: Neill, James O', et al.
Published: (2025)
by: Neill, James O', et al.
Published: (2025)
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
by: Sabry, Mohammed, et al.
Published: (2026)
by: Sabry, Mohammed, et al.
Published: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
AQUA: A Large Language Model for Aquaculture & Fisheries
by: Narisetty, Praneeth, et al.
Published: (2025)
by: Narisetty, Praneeth, et al.
Published: (2025)
Generalized Adaptive Transfer Network: Enhancing Transfer Learning in Reinforcement Learning Across Domains
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
The Impact of Inference Acceleration on Bias of LLMs
by: Kirsten, Elisabeth, et al.
Published: (2024)
by: Kirsten, Elisabeth, et al.
Published: (2024)
The Remarkable Robustness of LLMs: Stages of Inference?
by: Lad, Vedang, et al.
Published: (2024)
by: Lad, Vedang, et al.
Published: (2024)
Adaptive Soft Rolling KV Freeze with Entropy-Guided Recovery: Sublinear Memory Growth for Efficient LLM Inference
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
When Less is Enough: Efficient Inference via Collaborative Reasoning
by: Chen, Yilei, et al.
Published: (2026)
by: Chen, Yilei, et al.
Published: (2026)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
by: Krishnamurthy, Vikram
Published: (2026)
by: Krishnamurthy, Vikram
Published: (2026)
Hypertokens: Holographic Associative Memory in Tokenized LLMs
by: Augeri, Christopher James
Published: (2025)
by: Augeri, Christopher James
Published: (2025)
A Framework for Inference Inspired by Human Memory Mechanisms
by: Zeng, Xiangyu, et al.
Published: (2023)
by: Zeng, Xiangyu, et al.
Published: (2023)
Not All Layers of LLMs Are Necessary During Inference
by: Fan, Siqi, et al.
Published: (2024)
by: Fan, Siqi, et al.
Published: (2024)
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
by: Niu, Yue, et al.
Published: (2024)
by: Niu, Yue, et al.
Published: (2024)
Tracing Computation Density in LLMs
by: Kervadec, Corentin, et al.
Published: (2026)
by: Kervadec, Corentin, et al.
Published: (2026)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
Block-Attention for Efficient Prefilling
by: Ma, Dongyang, et al.
Published: (2024)
by: Ma, Dongyang, et al.
Published: (2024)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
by: Sharma, Agniv, et al.
Published: (2024)
by: Sharma, Agniv, et al.
Published: (2024)
Similar Items
-
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025) -
Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics
by: Acharjee, Jashaswimalya, et al.
Published: (2026) -
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
by: Yang, Dongquan, et al.
Published: (2025) -
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025) -
ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
by: Butler, Landon, et al.
Published: (2025)