EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Rajabzadeh, Hossein, Jafari, Aref, Sharma, Aman, Jami, Benyamin, Kwon, Hyock Ju, Ghodsi, Ali, Chen, Boxing, Rezagholizadeh, Mehdi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
di: Jamialahmadi, Benyamin, et al.
Pubblicazione: (2025)
di: Jamialahmadi, Benyamin, et al.
Pubblicazione: (2025)
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
di: Valipour, Mojtaba, et al.
Pubblicazione: (2023)
di: Valipour, Mojtaba, et al.
Pubblicazione: (2023)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
di: Sharma, Aman, et al.
Pubblicazione: (2025)
di: Sharma, Aman, et al.
Pubblicazione: (2025)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2024)
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2024)
Bayesian Mixture of Experts For Large Language Models
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2023)
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2023)
EMA-SAM: Exponential Moving-average for SAM-based PTMC Segmentation
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2026)
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2026)
DualSwinUnet++: An Enhanced Swin-Unet Architecture With Dual Decoders For PTMC Segmentation
di: Dialameh, Maryam, et al.
Pubblicazione: (2024)
di: Dialameh, Maryam, et al.
Pubblicazione: (2024)
SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
di: Jafari, Aref, et al.
Pubblicazione: (2025)
di: Jafari, Aref, et al.
Pubblicazione: (2025)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
di: Metel, Michael R., et al.
Pubblicazione: (2024)
di: Metel, Michael R., et al.
Pubblicazione: (2024)
On the importance of Data Scale in Pretraining Arabic Language Models
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
ReGLA: Refining Gated Linear Attention
di: Lu, Peng, et al.
Pubblicazione: (2025)
di: Lu, Peng, et al.
Pubblicazione: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
di: Metel, Michael R., et al.
Pubblicazione: (2024)
di: Metel, Michael R., et al.
Pubblicazione: (2024)
Metamaterial Bi-stable Vibration Absorbers for Railway Tracks: Experimental Study of Flexural Wave Control
di: Jafari, Ali, et al.
Pubblicazione: (2025)
di: Jafari, Ali, et al.
Pubblicazione: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
di: Guimarães, Heitor R., et al.
Pubblicazione: (2024)
di: Guimarães, Heitor R., et al.
Pubblicazione: (2024)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
di: Huang, Chenyang, et al.
Pubblicazione: (2024)
di: Huang, Chenyang, et al.
Pubblicazione: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
di: Mo, Fengran, et al.
Pubblicazione: (2024)
di: Mo, Fengran, et al.
Pubblicazione: (2024)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
di: Ghodsi, Ali
Pubblicazione: (2025)
di: Ghodsi, Ali
Pubblicazione: (2025)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
di: Fashi, Parsa Ashrafi, et al.
Pubblicazione: (2026)
di: Fashi, Parsa Ashrafi, et al.
Pubblicazione: (2026)
Self-similar group actions on ultragraphs and associated $C^*$-algebras
di: Larki, Hossein, et al.
Pubblicazione: (2025)
di: Larki, Hossein, et al.
Pubblicazione: (2025)
Minimality and effectiveness of the groupoid associated to a self-similar ultragraph
di: Larki, Hossein, et al.
Pubblicazione: (2025)
di: Larki, Hossein, et al.
Pubblicazione: (2025)
Depth and Autonomy: A Framework for Evaluating LLM Applications in Social Science Research
di: Sanaei, Ali, et al.
Pubblicazione: (2025)
di: Sanaei, Ali, et al.
Pubblicazione: (2025)
Zebra-Llama: Towards Extremely Efficient Hybrid Models
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
di: Wang, Suyuchen, et al.
Pubblicazione: (2024)
di: Wang, Suyuchen, et al.
Pubblicazione: (2024)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
di: Li, Guihong, et al.
Pubblicazione: (2025)
di: Li, Guihong, et al.
Pubblicazione: (2025)
Enhancing Predictive Accuracy in Pharmaceutical Sales Through An Ensemble Kernel Gaussian Process Regression Approach
di: Mirshekari, Shahin, et al.
Pubblicazione: (2024)
di: Mirshekari, Shahin, et al.
Pubblicazione: (2024)
Auto-Regressive Masked Diffusion Models
di: Karami, Mahdi, et al.
Pubblicazione: (2026)
di: Karami, Mahdi, et al.
Pubblicazione: (2026)
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
di: Karami, Mahdi, et al.
Pubblicazione: (2024)
di: Karami, Mahdi, et al.
Pubblicazione: (2024)
Att göra klass
di: Arping, Åsa
Pubblicazione: (2022)
di: Arping, Åsa
Pubblicazione: (2022)
Unbiased Regression-Adjusted Estimation of Average Treatment Effects in Randomized Controlled Trials
di: Abadie, Alberto, et al.
Pubblicazione: (2025)
di: Abadie, Alberto, et al.
Pubblicazione: (2025)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
di: Guan, Bryan, et al.
Pubblicazione: (2025)
di: Guan, Bryan, et al.
Pubblicazione: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
di: Wang, Xindi, et al.
Pubblicazione: (2024)
di: Wang, Xindi, et al.
Pubblicazione: (2024)
GraphPI: Efficient Protein Inference with Graph Neural Networks
di: Ma, Zheng, et al.
Pubblicazione: (2026)
di: Ma, Zheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024) -
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
di: Jamialahmadi, Benyamin, et al.
Pubblicazione: (2025) -
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
di: Valipour, Mojtaba, et al.
Pubblicazione: (2023) -
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
di: Sharma, Aman, et al.
Pubblicazione: (2025) -
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2024)