Latent Multi-Head Attention for Small Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mehta, Sushant, Dandekar, Raj, Dandekar, Rajat, Panat, Sreedath |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Decoders Laugh as Loud as Encoders
by: Borodach, Eli, et al.
Published: (2025)
by: Borodach, Eli, et al.
Published: (2025)
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024)
by: Shaikh, Ammar, et al.
Published: (2024)
Muon: Training and Trade-offs with Latent Attention and MoE
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Vision-Language Models display a strong gender bias
by: Konavoor, Aiswarya, et al.
Published: (2025)
by: Konavoor, Aiswarya, et al.
Published: (2025)
Simulating Misinformation Propagation in Social Networks using Large Language Models
by: Maurya, Raj Gaurav, et al.
Published: (2025)
by: Maurya, Raj Gaurav, et al.
Published: (2025)
NanoVLMs: How small can we go and still make coherent Vision Language Models?
by: Agarwalla, Mukund, et al.
Published: (2025)
by: Agarwalla, Mukund, et al.
Published: (2025)
Beyond Passive Viewing: A Pilot Study of a Hybrid Learning Platform Augmenting Video Lectures with Conversational AI
by: Abraar, Mohammed, et al.
Published: (2026)
by: Abraar, Mohammed, et al.
Published: (2026)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
by: Dawson, Fiifi, et al.
Published: (2024)
by: Dawson, Fiifi, et al.
Published: (2024)
Modeling chaotic Lorenz ODE System using Scientific Machine Learning
by: Kashyap, Sameera S, et al.
Published: (2024)
by: Kashyap, Sameera S, et al.
Published: (2024)
Regional Tiny Stories: Using Small Models to Compare Language Learning and Tokenizer Performance
by: Patil, Nirvan, et al.
Published: (2025)
by: Patil, Nirvan, et al.
Published: (2025)
Scientific machine learning in ecological systems: A study on the predator-prey dynamics
by: Devgupta, Ranabir, et al.
Published: (2024)
by: Devgupta, Ranabir, et al.
Published: (2024)
Forecasting N-Body Dynamics: A Comparative Study of Neural Ordinary Differential Equations and Universal Differential Equations
by: S, Suriya R, et al.
Published: (2025)
by: S, Suriya R, et al.
Published: (2025)
A comparative study of NeuralODE and Universal ODE approaches to solving Chandrasekhar White Dwarf equation
by: Martinez, Raymundo Vazquez, et al.
Published: (2024)
by: Martinez, Raymundo Vazquez, et al.
Published: (2024)
A Scientific Machine Learning Approach for Predicting and Forecasting Battery Degradation in Electric Vehicles
by: Murgai, Sharv, et al.
Published: (2024)
by: Murgai, Sharv, et al.
Published: (2024)
HULLMI: Human vs LLM identification with explainability
by: Joshi, Prathamesh Dinesh, et al.
Published: (2024)
by: Joshi, Prathamesh Dinesh, et al.
Published: (2024)
EARS-UDE: Evaluating Auditory Response in Sensory Overload with Universal Differential Equations
by: Salunke, Miheer, et al.
Published: (2025)
by: Salunke, Miheer, et al.
Published: (2025)
Adaptive tumor growth forecasting via neural & universal ODEs
by: Subramanian, Kavya, et al.
Published: (2025)
by: Subramanian, Kavya, et al.
Published: (2025)
BULL-ODE: Bullwhip Learning with Neural ODEs and Universal Differential Equations under Stochastic Demand
by: Naik, Nachiket N., et al.
Published: (2025)
by: Naik, Nachiket N., et al.
Published: (2025)
A study of Universal ODE approaches to predicting soil organic carbon
by: V. V, Satyanarayana Raju G., et al.
Published: (2025)
by: V. V, Satyanarayana Raju G., et al.
Published: (2025)
Three methods, one problem: Classical and AI approaches to no-three-in-line
by: Ramanathan, Pranav, et al.
Published: (2025)
by: Ramanathan, Pranav, et al.
Published: (2025)
Physics-Informed Neural ODEs with Scale-Aware Residuals for Learning Stiff Biophysical Dynamics
by: Kainth, Kamalpreet Singh, et al.
Published: (2025)
by: Kainth, Kamalpreet Singh, et al.
Published: (2025)
Understanding Malware Propagation Dynamics through Scientific Machine Learning
by: Pappu, Karthik, et al.
Published: (2025)
by: Pappu, Karthik, et al.
Published: (2025)
Physical Informed Neural Networks for modeling ocean pollutant
by: Battina, Karishma, et al.
Published: (2025)
by: Battina, Karishma, et al.
Published: (2025)
FASST: Fast LLM-based Simultaneous Speech Translation
by: Ouyang, Siqi, et al.
Published: (2024)
by: Ouyang, Siqi, et al.
Published: (2024)
Model-Grounded Symbolic Artificial Intelligence Systems Learning and Reasoning with Model-Grounded Symbolic Artificial Intelligence Systems
by: Chattopadhyay, Aniruddha, et al.
Published: (2025)
by: Chattopadhyay, Aniruddha, et al.
Published: (2025)
Analyzing Multi-Head Attention on Trojan BERT Models
by: Wang, Jingwei
Published: (2024)
by: Wang, Jingwei
Published: (2024)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
by: Ji, Tao, et al.
Published: (2025)
by: Ji, Tao, et al.
Published: (2025)
Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models
by: Ho, Zheng Yi, et al.
Published: (2024)
by: Ho, Zheng Yi, et al.
Published: (2024)
Exploring the Robustness of Language Models for Tabular Question Answering via Attention Analysis
by: Bhandari, Kushal Raj, et al.
Published: (2024)
by: Bhandari, Kushal Raj, et al.
Published: (2024)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models
by: Bose, Kartik, et al.
Published: (2026)
by: Bose, Kartik, et al.
Published: (2026)
On the Role of Attention Heads in Large Language Model Safety
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Reasoning or Overthinking: Evaluating Large Language Models on Financial Sentiment Analysis
by: Vamvourellis, Dimitris, et al.
Published: (2025)
by: Vamvourellis, Dimitris, et al.
Published: (2025)
Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
by: Tian, Yuxing, et al.
Published: (2026)
by: Tian, Yuxing, et al.
Published: (2026)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
by: Figliolia, Tomas, et al.
Published: (2025)
by: Figliolia, Tomas, et al.
Published: (2025)
Similar Items
-
Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
by: Mehta, Sushant, et al.
Published: (2025) -
Decoders Laugh as Loud as Encoders
by: Borodach, Eli, et al.
Published: (2025) -
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024) -
Muon: Training and Trade-offs with Latent Attention and MoE
by: Mehta, Sushant, et al.
Published: (2025) -
Vision-Language Models display a strong gender bias
by: Konavoor, Aiswarya, et al.
Published: (2025)