ReGLA: Refining Gated Linear Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Peng, Kobyzev, Ivan, Rezagholizadeh, Mehdi, Chen, Boxing, Langlais, Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
by: Lu, Peng, et al.
Published: (2023)
by: Lu, Peng, et al.
Published: (2023)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
ReGLA: Efficient Receptive-Field Modeling with Gated Linear Attention Network
by: Li, Junzhou, et al.
Published: (2026)
by: Li, Junzhou, et al.
Published: (2026)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
by: Huang, Chenyang, et al.
Published: (2024)
by: Huang, Chenyang, et al.
Published: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
by: Wang, Suyuchen, et al.
Published: (2024)
by: Wang, Suyuchen, et al.
Published: (2024)
Integral Transformer: Denoising Attention, Not Too Much Not Too Little
by: Kobyzev, Ivan, et al.
Published: (2025)
by: Kobyzev, Ivan, et al.
Published: (2025)
BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
by: Ghaddar, Abbas, et al.
Published: (2026)
by: Ghaddar, Abbas, et al.
Published: (2026)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
$\textit{BenchIE}^{FL}$ : A Manually Re-Annotated Fact-Based Open Information Extraction Benchmark
by: Lamarche, Fabrice, et al.
Published: (2024)
by: Lamarche, Fabrice, et al.
Published: (2024)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
by: Kavehzadeh, Parsa, et al.
Published: (2023)
by: Kavehzadeh, Parsa, et al.
Published: (2023)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
by: Kavehzadeh, Parsa, et al.
Published: (2024)
by: Kavehzadeh, Parsa, et al.
Published: (2024)
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Hyperparameter Optimization for Large Language Model Instruction-Tuning
by: Tribes, Christophe, et al.
Published: (2023)
by: Tribes, Christophe, et al.
Published: (2023)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025)
by: Li, Guihong, et al.
Published: (2025)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
by: Kobeissi, Amine, et al.
Published: (2026)
by: Kobeissi, Amine, et al.
Published: (2026)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
by: Jamialahmadi, Benyamin, et al.
Published: (2025)
by: Jamialahmadi, Benyamin, et al.
Published: (2025)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
by: Piedboeuf, Frédéric, et al.
Published: (2024)
by: Piedboeuf, Frédéric, et al.
Published: (2024)
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
by: Antoine, Elie, et al.
Published: (2024)
by: Antoine, Elie, et al.
Published: (2024)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
by: Guan, Bryan, et al.
Published: (2025)
by: Guan, Bryan, et al.
Published: (2025)
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Resona: Improving Context Copying in Linear Recurrence Models with Retrieval
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks
by: Antoine, Elie, et al.
Published: (2025)
by: Antoine, Elie, et al.
Published: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
Zebra-Llama: Towards Extremely Efficient Hybrid Models
by: Yang, Mingyu, et al.
Published: (2025)
by: Yang, Mingyu, et al.
Published: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt Optimization
by: Shi, Wenhang, et al.
Published: (2025)
by: Shi, Wenhang, et al.
Published: (2025)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
by: Chizhov, Pavel, et al.
Published: (2025)
by: Chizhov, Pavel, et al.
Published: (2025)
Toxicity of the Commons: Curating Open-Source Pre-Training Data
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
by: De, Soham, et al.
Published: (2024)
by: De, Soham, et al.
Published: (2024)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
by: Guo, Jialong, et al.
Published: (2024)
by: Guo, Jialong, et al.
Published: (2024)
EWEK-QA: Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering Systems
by: Dehghan, Mohammad, et al.
Published: (2024)
by: Dehghan, Mohammad, et al.
Published: (2024)
Similar Items
-
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
by: Lu, Peng, et al.
Published: (2023) -
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024) -
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
by: Metel, Michael R., et al.
Published: (2024) -
ReGLA: Efficient Receptive-Field Modeling with Gated Linear Attention Network
by: Li, Junzhou, et al.
Published: (2026) -
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
by: Huang, Chenyang, et al.
Published: (2024)