Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tyukin, Georgy, Dovonon, Gbetondji J-S, Kaddour, Jean, Minervini, Pasquale |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tensor Product Attention Is All You Need
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
by: Tyukin, Georgy
Published: (2024)
by: Tyukin, Georgy
Published: (2024)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
by: Gerber, Isaac
Published: (2025)
by: Gerber, Isaac
Published: (2025)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
by: Ding, Shiwei, et al.
Published: (2025)
by: Ding, Shiwei, et al.
Published: (2025)
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
by: Khan, Imran
Published: (2025)
by: Khan, Imran
Published: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
by: Nguyen-Tri, Quan, et al.
Published: (2025)
by: Nguyen-Tri, Quan, et al.
Published: (2025)
More Agents Is All You Need
by: Li, Junyou, et al.
Published: (2024)
by: Li, Junyou, et al.
Published: (2024)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
by: Gomaa, Eyad, et al.
Published: (2024)
by: Gomaa, Eyad, et al.
Published: (2024)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Block Rotation is All You Need for MXFP4 Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
by: Potamitis, Nearchos, et al.
Published: (2025)
by: Potamitis, Nearchos, et al.
Published: (2025)
Forget Attention: Importance-Aware Attention Is All You Need
by: Shin, Soohyeong, et al.
Published: (2026)
by: Shin, Soohyeong, et al.
Published: (2026)
You Need Better Attention Priors
by: Litman, Elon, et al.
Published: (2026)
by: Litman, Elon, et al.
Published: (2026)
Synthetic Data RL: Task Definition Is All You Need
by: Guo, Yiduo, et al.
Published: (2025)
by: Guo, Yiduo, et al.
Published: (2025)
SecEncoder: Logs are All You Need in Security
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
by: Guo, Zhiyu, et al.
Published: (2024)
by: Guo, Zhiyu, et al.
Published: (2024)
Attention is All You Need Until You Need Retention
by: Yaslioglu, M. Murat
Published: (2025)
by: Yaslioglu, M. Murat
Published: (2025)
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
Evidence Is All You Need: Ordering Imaging Studies via Language Model Alignment with the ACR Appropriateness Criteria
by: Yao, Michael S., et al.
Published: (2024)
by: Yao, Michael S., et al.
Published: (2024)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
by: Chan, Brian J, et al.
Published: (2024)
by: Chan, Brian J, et al.
Published: (2024)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
by: Yoon, Do-hyeon, et al.
Published: (2025)
by: Yoon, Do-hyeon, et al.
Published: (2025)
Memory Is All You Need: Testing How Model Memory Affects LLM Performance in Annotation Tasks
by: Timoneda, Joan C., et al.
Published: (2025)
by: Timoneda, Joan C., et al.
Published: (2025)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
by: Chen, Lingjiao, et al.
Published: (2024)
by: Chen, Lingjiao, et al.
Published: (2024)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
by: Steinmetz, Cody, et al.
Published: (2025)
by: Steinmetz, Cody, et al.
Published: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
by: Liu, Yiyang, et al.
Published: (2025)
by: Liu, Yiyang, et al.
Published: (2025)
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
by: Nehrdich, Sebastian, et al.
Published: (2024)
by: Nehrdich, Sebastian, et al.
Published: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
Geometry is All You Need: A Unified Taxonomy of Matrix and Tensor Factorization for Compression of Generative Language Models
by: Xu, Mingxue, et al.
Published: (2024)
by: Xu, Mingxue, et al.
Published: (2024)
Cooperation Is All You Need
by: Adeel, Ahsan, et al.
Published: (2023)
by: Adeel, Ahsan, et al.
Published: (2023)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
by: Gan, Chunjing, et al.
Published: (2024)
by: Gan, Chunjing, et al.
Published: (2024)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
All You Need is "Leet": Evading Hate-speech Detection AI
by: Kahu, Sampanna Yashwant, et al.
Published: (2025)
by: Kahu, Sampanna Yashwant, et al.
Published: (2025)
Some Attention is All You Need for Retrieval
by: Michalak, Felix, et al.
Published: (2025)
by: Michalak, Felix, et al.
Published: (2025)
Grimoire is All You Need for Enhancing Large Language Models
by: Chen, Ding, et al.
Published: (2024)
by: Chen, Ding, et al.
Published: (2024)
Context is All You Need
by: Delanois, Jean Erik, et al.
Published: (2026)
by: Delanois, Jean Erik, et al.
Published: (2026)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
by: Kulkarni, Atharva, et al.
Published: (2024)
by: Kulkarni, Atharva, et al.
Published: (2024)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Similar Items
-
Tensor Product Attention Is All You Need
by: Zhang, Yifan, et al.
Published: (2025) -
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026) -
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
by: Tyukin, Georgy
Published: (2024) -
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
by: Gerber, Isaac
Published: (2025) -
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
by: Ding, Shiwei, et al.
Published: (2025)