QCQA: Quality and Capacity-aware grouped Query Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Joshi, Vinay, Laddha, Prashant, Sinha, Shambhavi, Omer, Om Ji, Subramoney, Sreenivas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
von: Kundu, Souvik, et al.
Veröffentlicht: (2024)
von: Kundu, Souvik, et al.
Veröffentlicht: (2024)
Memorization in Attention-only Transformers
von: Dana, Léo, et al.
Veröffentlicht: (2024)
von: Dana, Léo, et al.
Veröffentlicht: (2024)
Weighted Grouped Query Attention in Transformers
von: Chinnakonduru, Sai Sena, et al.
Veröffentlicht: (2024)
von: Chinnakonduru, Sai Sena, et al.
Veröffentlicht: (2024)
Dynamic sparsity in tree-structured feed-forward layers at scale
von: Sedghi, Reza, et al.
Veröffentlicht: (2026)
von: Sedghi, Reza, et al.
Veröffentlicht: (2026)
SwiftMem: Fast Agentic Memory via Query-aware Indexing
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
von: Yu, Yijiong, et al.
Veröffentlicht: (2025)
von: Yu, Yijiong, et al.
Veröffentlicht: (2025)
Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models
von: Setty, Vinay
Veröffentlicht: (2024)
von: Setty, Vinay
Veröffentlicht: (2024)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
von: Joshi, Nisheeth, et al.
Veröffentlicht: (2025)
von: Joshi, Nisheeth, et al.
Veröffentlicht: (2025)
Empirical Capacity Model for Self-Attention Neural Networks
von: Härmä, Aki, et al.
Veröffentlicht: (2024)
von: Härmä, Aki, et al.
Veröffentlicht: (2024)
MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
LiveFC: A System for Live Fact-Checking of Audio Streams
von: V, Venktesh, et al.
Veröffentlicht: (2024)
von: V, Venktesh, et al.
Veröffentlicht: (2024)
Enhancing reasoning accuracy in large language models during inference time
von: Sharma, Vinay, et al.
Veröffentlicht: (2026)
von: Sharma, Vinay, et al.
Veröffentlicht: (2026)
Abstractive Text Summarization for Contemporary Sanskrit Prose: Issues and Challenges
von: Sinha, Shagun
Veröffentlicht: (2025)
von: Sinha, Shagun
Veröffentlicht: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
QueryNER: Segmentation of E-commerce Queries
von: Palen-Michel, Chester, et al.
Veröffentlicht: (2024)
von: Palen-Michel, Chester, et al.
Veröffentlicht: (2024)
OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models
von: Wan, Zhuoyue, et al.
Veröffentlicht: (2025)
von: Wan, Zhuoyue, et al.
Veröffentlicht: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Self-Attention Limits Working Memory Capacity of Transformer-Based Models
von: Gong, Dongyu, et al.
Veröffentlicht: (2024)
von: Gong, Dongyu, et al.
Veröffentlicht: (2024)
No Query, No Access
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
Generating Query-Focused Summarization Datasets from Query-Free Summarization Datasets
von: Chali, Yllias, et al.
Veröffentlicht: (2026)
von: Chali, Yllias, et al.
Veröffentlicht: (2026)
Compact Language Models via Pruning and Knowledge Distillation
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Skeletons Matter: Dynamic Data Augmentation for Text-to-Query
von: Ji, Yuchen, et al.
Veröffentlicht: (2025)
von: Ji, Yuchen, et al.
Veröffentlicht: (2025)
Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
von: Tian, Yuxing, et al.
Veröffentlicht: (2026)
von: Tian, Yuxing, et al.
Veröffentlicht: (2026)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
QueryPlot: Generating Geological Evidence Layers using Natural Language Queries for Mineral Exploration
von: Ye, Meng, et al.
Veröffentlicht: (2026)
von: Ye, Meng, et al.
Veröffentlicht: (2026)
No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
Uniform Discretized Integrated Gradients: An effective attribution based method for explaining large language models
von: Roy, Swarnava Sinha, et al.
Veröffentlicht: (2024)
von: Roy, Swarnava Sinha, et al.
Veröffentlicht: (2024)
Query-Efficient Planning with Language Models
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2024)
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2024)
TrICy: Trigger-guided Data-to-text Generation with Intent aware Attention-Copy
von: Agarwal, Vibhav, et al.
Veröffentlicht: (2024)
von: Agarwal, Vibhav, et al.
Veröffentlicht: (2024)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
von: Guan, Ziyi, et al.
Veröffentlicht: (2024)
von: Guan, Ziyi, et al.
Veröffentlicht: (2024)
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
von: Serai, Prashant, et al.
Veröffentlicht: (2024)
von: Serai, Prashant, et al.
Veröffentlicht: (2024)
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
von: V, Venktesh, et al.
Veröffentlicht: (2024)
von: V, Venktesh, et al.
Veröffentlicht: (2024)
CIE: Controlling Language Model Text Generations Using Continuous Signals
von: Samuel, Vinay, et al.
Veröffentlicht: (2025)
von: Samuel, Vinay, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
von: Kundu, Souvik, et al.
Veröffentlicht: (2024) -
Memorization in Attention-only Transformers
von: Dana, Léo, et al.
Veröffentlicht: (2024) -
Weighted Grouped Query Attention in Transformers
von: Chinnakonduru, Sai Sena, et al.
Veröffentlicht: (2024) -
Dynamic sparsity in tree-structured feed-forward layers at scale
von: Sedghi, Reza, et al.
Veröffentlicht: (2026) -
SwiftMem: Fast Agentic Memory via Query-aware Indexing
von: Tian, Anxin, et al.
Veröffentlicht: (2026)