LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Zikai, Zhang, Qizheng, Kumbong, Hermann, Olukotun, Kunle |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
von: Du, Dayou, et al.
Veröffentlicht: (2025)
von: Du, Dayou, et al.
Veröffentlicht: (2025)
LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
von: Ko, Sho, et al.
Veröffentlicht: (2025)
von: Ko, Sho, et al.
Veröffentlicht: (2025)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
von: Lacouture, Rubens, et al.
Veröffentlicht: (2025)
von: Lacouture, Rubens, et al.
Veröffentlicht: (2025)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
A$^3$PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
von: Sohn, Gina, et al.
Veröffentlicht: (2025)
von: Sohn, Gina, et al.
Veröffentlicht: (2025)
Latency Based Tiling
von: Cashman, Jack
Veröffentlicht: (2025)
von: Cashman, Jack
Veröffentlicht: (2025)
HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
Examem: Low-Overhead Memory Instrumentation for Intelligent Memory Systems
von: Poduval, Ashwin, et al.
Veröffentlicht: (2024)
von: Poduval, Ashwin, et al.
Veröffentlicht: (2024)
A2Q+: Improving Accumulator-Aware Weight Quantization
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
von: Singh, Raul, et al.
Veröffentlicht: (2025)
von: Singh, Raul, et al.
Veröffentlicht: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
GeneTEK: Low-power, high-performance and scalable FPGA architecture for exact unit-cost edit distance
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
von: Chen, Yao, et al.
Veröffentlicht: (2022)
von: Chen, Yao, et al.
Veröffentlicht: (2022)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
Search Your Block Floating Point Scales!
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
Graph neural networks with configuration cross-attention for tensor compilers
von: Khizbullin, Dmitrii, et al.
Veröffentlicht: (2024)
von: Khizbullin, Dmitrii, et al.
Veröffentlicht: (2024)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
von: Zuepke, Alexander, et al.
Veröffentlicht: (2026)
von: Zuepke, Alexander, et al.
Veröffentlicht: (2026)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
von: Mao, Shunyu, et al.
Veröffentlicht: (2024)
von: Mao, Shunyu, et al.
Veröffentlicht: (2024)
Cleaning up the Mess: Re-Evaluating the Real-System Modeling Accuracy of Ramulator 2.0
von: Bostanci, F. Nisa, et al.
Veröffentlicht: (2025)
von: Bostanci, F. Nisa, et al.
Veröffentlicht: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
von: Vahdatniya, Parmida, et al.
Veröffentlicht: (2025)
von: Vahdatniya, Parmida, et al.
Veröffentlicht: (2025)
LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and Optimization
von: Sarkar, Rishov, et al.
Veröffentlicht: (2024)
von: Sarkar, Rishov, et al.
Veröffentlicht: (2024)
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
von: Ko, Sho, et al.
Veröffentlicht: (2024)
von: Ko, Sho, et al.
Veröffentlicht: (2024)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
von: Xie, Xilong, et al.
Veröffentlicht: (2025)
von: Xie, Xilong, et al.
Veröffentlicht: (2025)
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025) -
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
von: Du, Dayou, et al.
Veröffentlicht: (2025) -
LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025) -
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
von: Ko, Sho, et al.
Veröffentlicht: (2025) -
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)