QCQA: Quality and Capacity-aware grouped Query Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Joshi, Vinay, Laddha, Prashant, Sinha, Shambhavi, Omer, Om Ji, Subramoney, Sreenivas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
by: Kundu, Souvik, et al.
Published: (2024)
by: Kundu, Souvik, et al.
Published: (2024)
Memorization in Attention-only Transformers
by: Dana, Léo, et al.
Published: (2024)
by: Dana, Léo, et al.
Published: (2024)
Weighted Grouped Query Attention in Transformers
by: Chinnakonduru, Sai Sena, et al.
Published: (2024)
by: Chinnakonduru, Sai Sena, et al.
Published: (2024)
Dynamic sparsity in tree-structured feed-forward layers at scale
by: Sedghi, Reza, et al.
Published: (2026)
by: Sedghi, Reza, et al.
Published: (2026)
SwiftMem: Fast Agentic Memory via Query-aware Indexing
by: Tian, Anxin, et al.
Published: (2026)
by: Tian, Anxin, et al.
Published: (2026)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
by: Devoto, Alessio, et al.
Published: (2025)
by: Devoto, Alessio, et al.
Published: (2025)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models
by: Setty, Vinay
Published: (2024)
by: Setty, Vinay
Published: (2024)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
by: Xiao, Qingfa, et al.
Published: (2025)
by: Xiao, Qingfa, et al.
Published: (2025)
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
by: Joshi, Nisheeth, et al.
Published: (2025)
by: Joshi, Nisheeth, et al.
Published: (2025)
Empirical Capacity Model for Self-Attention Neural Networks
by: Härmä, Aki, et al.
Published: (2024)
by: Härmä, Aki, et al.
Published: (2024)
MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries
by: Ghosh, Akash, et al.
Published: (2024)
by: Ghosh, Akash, et al.
Published: (2024)
LiveFC: A System for Live Fact-Checking of Audio Streams
by: V, Venktesh, et al.
Published: (2024)
by: V, Venktesh, et al.
Published: (2024)
Enhancing reasoning accuracy in large language models during inference time
by: Sharma, Vinay, et al.
Published: (2026)
by: Sharma, Vinay, et al.
Published: (2026)
Abstractive Text Summarization for Contemporary Sanskrit Prose: Issues and Challenges
by: Sinha, Shagun
Published: (2025)
by: Sinha, Shagun
Published: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
by: Chen, Yingfa, et al.
Published: (2025)
by: Chen, Yingfa, et al.
Published: (2025)
QueryNER: Segmentation of E-commerce Queries
by: Palen-Michel, Chester, et al.
Published: (2024)
by: Palen-Michel, Chester, et al.
Published: (2024)
OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models
by: Wan, Zhuoyue, et al.
Published: (2025)
by: Wan, Zhuoyue, et al.
Published: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
Cross-Attention Speculative Decoding
by: Zhong, Wei, et al.
Published: (2025)
by: Zhong, Wei, et al.
Published: (2025)
Self-Attention Limits Working Memory Capacity of Transformer-Based Models
by: Gong, Dongyu, et al.
Published: (2024)
by: Gong, Dongyu, et al.
Published: (2024)
No Query, No Access
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
Generating Query-Focused Summarization Datasets from Query-Free Summarization Datasets
by: Chali, Yllias, et al.
Published: (2026)
by: Chali, Yllias, et al.
Published: (2026)
Compact Language Models via Pruning and Knowledge Distillation
by: Muralidharan, Saurav, et al.
Published: (2024)
by: Muralidharan, Saurav, et al.
Published: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Skeletons Matter: Dynamic Data Augmentation for Text-to-Query
by: Ji, Yuchen, et al.
Published: (2025)
by: Ji, Yuchen, et al.
Published: (2025)
Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
by: Tian, Yuxing, et al.
Published: (2026)
by: Tian, Yuxing, et al.
Published: (2026)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
QueryPlot: Generating Geological Evidence Layers using Natural Language Queries for Mineral Exploration
by: Ye, Meng, et al.
Published: (2026)
by: Ye, Meng, et al.
Published: (2026)
No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
by: Rauch, Lukas, et al.
Published: (2025)
by: Rauch, Lukas, et al.
Published: (2025)
Uniform Discretized Integrated Gradients: An effective attribution based method for explaining large language models
by: Roy, Swarnava Sinha, et al.
Published: (2024)
by: Roy, Swarnava Sinha, et al.
Published: (2024)
Query-Efficient Planning with Language Models
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2024)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2024)
TrICy: Trigger-guided Data-to-text Generation with Intent aware Attention-Copy
by: Agarwal, Vibhav, et al.
Published: (2024)
by: Agarwal, Vibhav, et al.
Published: (2024)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
by: Serai, Prashant, et al.
Published: (2024)
by: Serai, Prashant, et al.
Published: (2024)
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
by: V, Venktesh, et al.
Published: (2024)
by: V, Venktesh, et al.
Published: (2024)
CIE: Controlling Language Model Text Generations Using Continuous Signals
by: Samuel, Vinay, et al.
Published: (2025)
by: Samuel, Vinay, et al.
Published: (2025)
Similar Items
-
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
by: Kundu, Souvik, et al.
Published: (2024) -
Memorization in Attention-only Transformers
by: Dana, Léo, et al.
Published: (2024) -
Weighted Grouped Query Attention in Transformers
by: Chinnakonduru, Sai Sena, et al.
Published: (2024) -
Dynamic sparsity in tree-structured feed-forward layers at scale
by: Sedghi, Reza, et al.
Published: (2026) -
SwiftMem: Fast Agentic Memory via Query-aware Indexing
by: Tian, Anxin, et al.
Published: (2026)