Salvato in:
| Autori principali: | Zhang, Cheng, Cheng, Jianyi, Constantinides, George A., Zhao, Yiren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.02446 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unlocking the Global Synergies in Low-Rank Adapters
di: Zhang, Zixi, et al.
Pubblicazione: (2024)
di: Zhang, Zixi, et al.
Pubblicazione: (2024)
QERA: an Analytical Framework for Quantization Error Reconstruction
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
A3 : an Analytical Low-Rank Approximation Framework for Attention
di: Wong, Jeffrey T. H., et al.
Pubblicazione: (2025)
di: Wong, Jeffrey T. H., et al.
Pubblicazione: (2025)
Scaling Laws For Mixed Quantization
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
di: Zhang, Cheng, et al.
Pubblicazione: (2023)
di: Zhang, Cheng, et al.
Pubblicazione: (2023)
Low-Rank Quantization-Aware Training for LLMs
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
LoQT: Low-Rank Adapters for Quantized Pretraining
di: Loeschcke, Sebastian, et al.
Pubblicazione: (2024)
di: Loeschcke, Sebastian, et al.
Pubblicazione: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
di: Cai, Wen-Pu, et al.
Pubblicazione: (2024)
di: Cai, Wen-Pu, et al.
Pubblicazione: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
di: Zhang, Rongzhi, et al.
Pubblicazione: (2024)
di: Zhang, Rongzhi, et al.
Pubblicazione: (2024)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
di: Song, Guanghui, et al.
Pubblicazione: (2025)
di: Song, Guanghui, et al.
Pubblicazione: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
di: Cheng, Wenhua, et al.
Pubblicazione: (2023)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
di: Xiong, Boya, et al.
Pubblicazione: (2025)
di: Xiong, Boya, et al.
Pubblicazione: (2025)
On the Existence and Behavior of Secondary Attention Sinks
di: Wong, Jeffrey T. H., et al.
Pubblicazione: (2025)
di: Wong, Jeffrey T. H., et al.
Pubblicazione: (2025)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
LoRMA: Low-Rank Multiplicative Adaptation for LLMs
di: Bihany, Harsh, et al.
Pubblicazione: (2025)
di: Bihany, Harsh, et al.
Pubblicazione: (2025)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
di: Sengupta, Ayan, et al.
Pubblicazione: (2024)
di: Sengupta, Ayan, et al.
Pubblicazione: (2024)
Optimised Grouped-Query Attention Mechanism for Transformers
di: Chen, Yuang, et al.
Pubblicazione: (2024)
di: Chen, Yuang, et al.
Pubblicazione: (2024)
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2023)
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2023)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
di: Cho, Yoonjun, et al.
Pubblicazione: (2025)
di: Cho, Yoonjun, et al.
Pubblicazione: (2025)
An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
di: Dadgarnia, Alireza, et al.
Pubblicazione: (2026)
di: Dadgarnia, Alireza, et al.
Pubblicazione: (2026)
Understanding and Mitigating Errors of LLM-Generated RTL Code
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
di: Huang, Xijie, et al.
Pubblicazione: (2024)
di: Huang, Xijie, et al.
Pubblicazione: (2024)
FrameQuant: Flexible Low-Bit Quantization for Transformers
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
di: Lee, Changhun, et al.
Pubblicazione: (2024)
di: Lee, Changhun, et al.
Pubblicazione: (2024)
How Does Quantization Affect Multilingual LLMs?
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
di: Abitante, João Vitor Boer, et al.
Pubblicazione: (2026)
di: Abitante, João Vitor Boer, et al.
Pubblicazione: (2026)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
di: Cheng, Yun, et al.
Pubblicazione: (2026)
di: Cheng, Yun, et al.
Pubblicazione: (2026)
BiSup: Bidirectional Quantization Error Suppression for Large Language Models
di: Zou, Minghui, et al.
Pubblicazione: (2024)
di: Zou, Minghui, et al.
Pubblicazione: (2024)
Optimizing Large Language Model Training Using FP4 Quantization
di: Wang, Ruizhe, et al.
Pubblicazione: (2025)
di: Wang, Ruizhe, et al.
Pubblicazione: (2025)
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
di: Wang, Xinyi, et al.
Pubblicazione: (2025)
di: Wang, Xinyi, et al.
Pubblicazione: (2025)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
di: Tang, Chuanyu, et al.
Pubblicazione: (2024)
di: Tang, Chuanyu, et al.
Pubblicazione: (2024)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
di: Liao, Xutao, et al.
Pubblicazione: (2024)
di: Liao, Xutao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Unlocking the Global Synergies in Low-Rank Adapters
di: Zhang, Zixi, et al.
Pubblicazione: (2024) -
QERA: an Analytical Framework for Quantization Error Reconstruction
di: Zhang, Cheng, et al.
Pubblicazione: (2024) -
A3 : an Analytical Low-Rank Approximation Framework for Attention
di: Wong, Jeffrey T. H., et al.
Pubblicazione: (2025) -
Scaling Laws For Mixed Quantization
di: Cao, Zeyu, et al.
Pubblicazione: (2024) -
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
di: Zhang, Cheng, et al.
Pubblicazione: (2023)