Inference-Friendly Models With MixAttention
Fuente:
arXiv
Saved in:
| Main Authors: | Rajput, Shashank, Sheng, Ying, Owen, Sean, Chiley, Vitaliy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse Upcycling: Inference Inefficient Finetuning
by: Doubov, Sasha, et al.
Published: (2024)
by: Doubov, Sasha, et al.
Published: (2024)
LoRA Learns Less and Forgets Less
by: Biderman, Dan, et al.
Published: (2024)
by: Biderman, Dan, et al.
Published: (2024)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
by: Merrill, Scott, et al.
Published: (2026)
by: Merrill, Scott, et al.
Published: (2026)
Self-Selected Attention Span for Accelerating Large Language Model Inference
by: Jin, Tian, et al.
Published: (2024)
by: Jin, Tian, et al.
Published: (2024)
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
by: Lee, Sangyun, et al.
Published: (2026)
by: Lee, Sangyun, et al.
Published: (2026)
CATP: Cross-Attention Token Pruning for Accuracy Preserved Multimodal Model Inference
by: Liao, Ruqi, et al.
Published: (2024)
by: Liao, Ruqi, et al.
Published: (2024)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
by: Zhang, Hanzhi, et al.
Published: (2025)
by: Zhang, Hanzhi, et al.
Published: (2025)
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
by: Tang, Yaohua, et al.
Published: (2025)
by: Tang, Yaohua, et al.
Published: (2025)
Optimizing Temperature for Language Models with Multi-Sample Inference
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
by: Yona, Itay, et al.
Published: (2026)
by: Yona, Itay, et al.
Published: (2026)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
by: Gim, In, et al.
Published: (2023)
by: Gim, In, et al.
Published: (2023)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
by: Gu, Yifeng, et al.
Published: (2025)
by: Gu, Yifeng, et al.
Published: (2025)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
AutoMix: Automatically Mixing Language Models
by: Aggarwal, Pranjal, et al.
Published: (2023)
by: Aggarwal, Pranjal, et al.
Published: (2023)
How Open Must Language Models be to Enable Reliable Scientific Inference?
by: Michaelov, James A., et al.
Published: (2026)
by: Michaelov, James A., et al.
Published: (2026)
Attention Basin: Why Contextual Position Matters in Large Language Models
by: Yi, Zihao, et al.
Published: (2025)
by: Yi, Zihao, et al.
Published: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Positional Encoding via Token-Aware Phase Attention
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
by: Jadon, Aryan, et al.
Published: (2025)
by: Jadon, Aryan, et al.
Published: (2025)
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models
by: Kaiser, Caspar, et al.
Published: (2026)
by: Kaiser, Caspar, et al.
Published: (2026)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
by: Begin, James, et al.
Published: (2025)
by: Begin, James, et al.
Published: (2025)
PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead
by: Tan, Tao, et al.
Published: (2024)
by: Tan, Tao, et al.
Published: (2024)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
by: Ji, Tao, et al.
Published: (2025)
by: Ji, Tao, et al.
Published: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Toward Sustainable GenAI using Generation Directives for Carbon-Friendly Large Language Model Inference
by: Li, Baolin, et al.
Published: (2024)
by: Li, Baolin, et al.
Published: (2024)
dInfer: An Efficient Inference Framework for Diffusion Language Models
by: Ma, Yuxin, et al.
Published: (2025)
by: Ma, Yuxin, et al.
Published: (2025)
Attention Sinks in Diffusion Language Models
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
COMET: Generating Commit Messages using Delta Graph Context Representation
by: Mandli, Abhinav Reddy, et al.
Published: (2024)
by: Mandli, Abhinav Reddy, et al.
Published: (2024)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
by: Chernyshev, Konstantin, et al.
Published: (2024)
by: Chernyshev, Konstantin, et al.
Published: (2024)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
ASRJam: Human-Friendly AI Speech Jamming to Prevent Automated Phone Scams
by: Grabovski, Freddie, et al.
Published: (2025)
by: Grabovski, Freddie, et al.
Published: (2025)
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Cross-Attention Watermarking of Large Language Models
by: Baldassini, Folco Bertini, et al.
Published: (2024)
by: Baldassini, Folco Bertini, et al.
Published: (2024)
Similar Items
-
Sparse Upcycling: Inference Inefficient Finetuning
by: Doubov, Sasha, et al.
Published: (2024) -
LoRA Learns Less and Forgets Less
by: Biderman, Dan, et al.
Published: (2024) -
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
by: Merrill, Scott, et al.
Published: (2026) -
Self-Selected Attention Span for Accelerating Large Language Model Inference
by: Jin, Tian, et al.
Published: (2024) -
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
by: Lee, Sangyun, et al.
Published: (2026)