Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Tyukin, Georgy, Dovonon, Gbetondji J-S, Kaddour, Jean, Minervini, Pasquale |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
di: Tyukin, Georgy
Pubblicazione: (2024)
di: Tyukin, Georgy
Pubblicazione: (2024)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
di: Gerber, Isaac
Pubblicazione: (2025)
di: Gerber, Isaac
Pubblicazione: (2025)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
di: Ding, Shiwei, et al.
Pubblicazione: (2025)
di: Ding, Shiwei, et al.
Pubblicazione: (2025)
Agentic Uncertainty Reveals Agentic Overconfidence
di: Kaddour, Jean, et al.
Pubblicazione: (2026)
di: Kaddour, Jean, et al.
Pubblicazione: (2026)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
di: Khan, Imran
Pubblicazione: (2025)
di: Khan, Imran
Pubblicazione: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
More Agents Is All You Need
di: Li, Junyou, et al.
Pubblicazione: (2024)
di: Li, Junyou, et al.
Pubblicazione: (2024)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
di: Gomaa, Eyad, et al.
Pubblicazione: (2024)
di: Gomaa, Eyad, et al.
Pubblicazione: (2024)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
di: Li, Pengyi, et al.
Pubblicazione: (2025)
di: Li, Pengyi, et al.
Pubblicazione: (2025)
Block Rotation is All You Need for MXFP4 Quantization
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
di: Potamitis, Nearchos, et al.
Pubblicazione: (2025)
di: Potamitis, Nearchos, et al.
Pubblicazione: (2025)
Forget Attention: Importance-Aware Attention Is All You Need
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
You Need Better Attention Priors
di: Litman, Elon, et al.
Pubblicazione: (2026)
di: Litman, Elon, et al.
Pubblicazione: (2026)
Synthetic Data RL: Task Definition Is All You Need
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
SecEncoder: Logs are All You Need in Security
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
di: Guo, Zhiyu, et al.
Pubblicazione: (2024)
di: Guo, Zhiyu, et al.
Pubblicazione: (2024)
Attention is All You Need Until You Need Retention
di: Yaslioglu, M. Murat
Pubblicazione: (2025)
di: Yaslioglu, M. Murat
Pubblicazione: (2025)
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
di: Jiang, Yi, et al.
Pubblicazione: (2025)
di: Jiang, Yi, et al.
Pubblicazione: (2025)
Evidence Is All You Need: Ordering Imaging Studies via Language Model Alignment with the ACR Appropriateness Criteria
di: Yao, Michael S., et al.
Pubblicazione: (2024)
di: Yao, Michael S., et al.
Pubblicazione: (2024)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
di: Chan, Brian J, et al.
Pubblicazione: (2024)
di: Chan, Brian J, et al.
Pubblicazione: (2024)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
Memory Is All You Need: Testing How Model Memory Affects LLM Performance in Annotation Tasks
di: Timoneda, Joan C., et al.
Pubblicazione: (2025)
di: Timoneda, Joan C., et al.
Pubblicazione: (2025)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
di: Steinmetz, Cody, et al.
Pubblicazione: (2025)
di: Steinmetz, Cody, et al.
Pubblicazione: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
di: Nehrdich, Sebastian, et al.
Pubblicazione: (2024)
di: Nehrdich, Sebastian, et al.
Pubblicazione: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
di: Hernandez, Adriano
Pubblicazione: (2024)
di: Hernandez, Adriano
Pubblicazione: (2024)
Geometry is All You Need: A Unified Taxonomy of Matrix and Tensor Factorization for Compression of Generative Language Models
di: Xu, Mingxue, et al.
Pubblicazione: (2024)
di: Xu, Mingxue, et al.
Pubblicazione: (2024)
Cooperation Is All You Need
di: Adeel, Ahsan, et al.
Pubblicazione: (2023)
di: Adeel, Ahsan, et al.
Pubblicazione: (2023)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
di: Gan, Chunjing, et al.
Pubblicazione: (2024)
di: Gan, Chunjing, et al.
Pubblicazione: (2024)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
di: LeVine, Will, et al.
Pubblicazione: (2025)
di: LeVine, Will, et al.
Pubblicazione: (2025)
All You Need is "Leet": Evading Hate-speech Detection AI
di: Kahu, Sampanna Yashwant, et al.
Pubblicazione: (2025)
di: Kahu, Sampanna Yashwant, et al.
Pubblicazione: (2025)
Some Attention is All You Need for Retrieval
di: Michalak, Felix, et al.
Pubblicazione: (2025)
di: Michalak, Felix, et al.
Pubblicazione: (2025)
Grimoire is All You Need for Enhancing Large Language Models
di: Chen, Ding, et al.
Pubblicazione: (2024)
di: Chen, Ding, et al.
Pubblicazione: (2024)
Context is All You Need
di: Delanois, Jean Erik, et al.
Pubblicazione: (2026)
di: Delanois, Jean Erik, et al.
Pubblicazione: (2026)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
What Matters in Transformers? Not All Attention is Needed
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025) -
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026) -
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
di: Tyukin, Georgy
Pubblicazione: (2024) -
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
di: Gerber, Isaac
Pubblicazione: (2025) -
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
di: Ding, Shiwei, et al.
Pubblicazione: (2025)