Saved in:
| Main Authors: | Jin, Qingyun, Song, Xiaohui, Zhou, Feng, Qin, Zengchang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.20677 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SyntheT2C: Generating Synthetic Data for Fine-Tuning Large Language Models on the Text2Cypher Task
by: Zhong, Ziije, et al.
Published: (2024)
by: Zhong, Ziije, et al.
Published: (2024)
Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
FlashMem: Distilling Intrinsic Latent Memory via Computation Reuse
by: Hou, Yubo, et al.
Published: (2026)
by: Hou, Yubo, et al.
Published: (2026)
Sensitivity-Positional Co-Localization in GQA Transformers
by: Rao, Manoj Chandrashekar
Published: (2026)
by: Rao, Manoj Chandrashekar
Published: (2026)
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation
by: Qu, Ge, et al.
Published: (2024)
by: Qu, Ge, et al.
Published: (2024)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
An Innovative CGL-MHA Model for Sarcasm Sentiment Recognition Using the MindSpore Framework
by: Qin, Zhenkai, et al.
Published: (2024)
by: Qin, Zhenkai, et al.
Published: (2024)
CADReN: Contextual Anchor-Driven Relational Network for Controllable Cross-Graphs Node Importance Estimation
by: Zhong, Zijie, et al.
Published: (2024)
by: Zhong, Zijie, et al.
Published: (2024)
Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation
by: Zhong, Zijie, et al.
Published: (2024)
by: Zhong, Zijie, et al.
Published: (2024)
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
by: Khalifa, Muhammad, et al.
Published: (2024)
by: Khalifa, Muhammad, et al.
Published: (2024)
Architecture-Dependent Processing Mode Dynamics in Transformer Attention: Opposing Transitions in MHA, GQA, and MoE Models
by: Ichikawa, Yuki
Published: (2026)
by: Ichikawa, Yuki
Published: (2026)
Knocking-Heads Attention
by: Zhou, Zhanchao, et al.
Published: (2025)
by: Zhou, Zhanchao, et al.
Published: (2025)
N2N-GQA: Noise-to-Narrative for Graph-Based Table-Text Question Answering Using LLMs
by: Sharafath, Mohamed, et al.
Published: (2026)
by: Sharafath, Mohamed, et al.
Published: (2026)
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
by: Li, Minchong, et al.
Published: (2024)
by: Li, Minchong, et al.
Published: (2024)
Conversations: Love Them, Hate Them, Steer Them
by: Chebrolu, Niranjan, et al.
Published: (2025)
by: Chebrolu, Niranjan, et al.
Published: (2025)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
by: Lin, Xihui, et al.
Published: (2024)
by: Lin, Xihui, et al.
Published: (2024)
CALM: Unleashing the Cross-Lingual Self-Aligning Ability of Language Model Question Answering
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
R2GQA: Retriever-Reader-Generator Question Answering System to Support Students Understanding Legal Regulations in Higher Education
by: Do, Phuc-Tinh Pham, et al.
Published: (2024)
by: Do, Phuc-Tinh Pham, et al.
Published: (2024)
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging
by: Gain, Baban, et al.
Published: (2026)
by: Gain, Baban, et al.
Published: (2026)
Attention Heads of Large Language Models: A Survey
by: Zheng, Zifan, et al.
Published: (2024)
by: Zheng, Zifan, et al.
Published: (2024)
Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
by: Liu, Jiale, et al.
Published: (2025)
by: Liu, Jiale, et al.
Published: (2025)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
Split and Merge: Aligning Position Biases in LLM-based Evaluators
by: Li, Zongjie, et al.
Published: (2023)
by: Li, Zongjie, et al.
Published: (2023)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
ACE-Merging: Data-Free Model Merging with Adaptive Covariance Estimation
by: Xu, Bo, et al.
Published: (2026)
by: Xu, Bo, et al.
Published: (2026)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
Effectively Compress KV Heads for LLM
by: Yu, Hao, et al.
Published: (2024)
by: Yu, Hao, et al.
Published: (2024)
Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
FroM: Frobenius Norm-Based Data-Free Adaptive Model Merging
by: Li, Zijian, et al.
Published: (2025)
by: Li, Zijian, et al.
Published: (2025)
Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs
by: Zheng, Chen, et al.
Published: (2024)
by: Zheng, Chen, et al.
Published: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2025)
by: Feng, Yujie, et al.
Published: (2025)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
by: Lin, Xiaolin, et al.
Published: (2025)
by: Lin, Xiaolin, et al.
Published: (2025)
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)
by: Elhelo, Amit, et al.
Published: (2024)
Interpreting Transformers Through Attention Head Intervention
by: Kadem, Mason, et al.
Published: (2026)
by: Kadem, Mason, et al.
Published: (2026)
MossNet: Mixture of State-Space Experts is a Multi-Head Attention
by: Tuli, Shikhar, et al.
Published: (2025)
by: Tuli, Shikhar, et al.
Published: (2025)
LLMs Can Plan Only If We Tell Them
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
1bit-Merging: Dynamic Quantized Merging for Large Language Models
by: Liu, Shuqi, et al.
Published: (2025)
by: Liu, Shuqi, et al.
Published: (2025)
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
by: Haresamudram, Harish, et al.
Published: (2024)
by: Haresamudram, Harish, et al.
Published: (2024)
Similar Items
-
SyntheT2C: Generating Synthetic Data for Fine-Tuning Large Language Models on the Text2Cypher Task
by: Zhong, Ziije, et al.
Published: (2024) -
Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
by: Gumaan, Esmail
Published: (2025) -
FlashMem: Distilling Intrinsic Latent Memory via Computation Reuse
by: Hou, Yubo, et al.
Published: (2026) -
Sensitivity-Positional Co-Localization in GQA Transformers
by: Rao, Manoj Chandrashekar
Published: (2026) -
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation
by: Qu, Ge, et al.
Published: (2024)