HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Taiqiang, Ding, Chenchen, Zhou, Wenyong, Cheng, Yuxin, Feng, Xincheng, Wang, Shuqi, Xu, Wendong, Shi, Chufan, Liu, Zhengwu, Wong, Ngai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
In-Memory Computing Architecture for Efficient Hardware Security
von: Ajmi, Hala, et al.
Veröffentlicht: (2024)
von: Ajmi, Hala, et al.
Veröffentlicht: (2024)
Accelerating PageRank Algorithmic Tasks with a new Programmable Hardware Architecture
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2024)
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
Towards the Certification of Hybrid Architectures: Analysing Interference on Hardware Accelerators through PML
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
Hardware Memory Management for Future Mobile Hybrid Memory Systems
von: Wen, Fei, et al.
Veröffentlicht: (2020)
von: Wen, Fei, et al.
Veröffentlicht: (2020)
PUMA: Efficient and Low-Cost Memory Allocation and Alignment Support for Processing-Using-Memory Architectures
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
QED: Scalable Verification of Hardware Memory Consistency
von: Ravi, Gokulan, et al.
Veröffentlicht: (2024)
von: Ravi, Gokulan, et al.
Veröffentlicht: (2024)
An Event-Driven Spiking Compute-In-Memory Macro based on SOT-MRAM
von: Yu, Deyang, et al.
Veröffentlicht: (2025)
von: Yu, Deyang, et al.
Veröffentlicht: (2025)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
von: Hwang, Soojin, et al.
Veröffentlicht: (2025)
von: Hwang, Soojin, et al.
Veröffentlicht: (2025)
Xpikeformer: Hybrid Analog-Digital Hardware Acceleration for Spiking Transformers
von: Song, Zihang, et al.
Veröffentlicht: (2024)
von: Song, Zihang, et al.
Veröffentlicht: (2024)
Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient Redistribution
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
von: Kundu, Souvik, et al.
Veröffentlicht: (2024)
von: Kundu, Souvik, et al.
Veröffentlicht: (2024)
HLStrans: Dataset for C-to-HLS Hardware Code Synthesis
von: Zou, Qingyun, et al.
Veröffentlicht: (2025)
von: Zou, Qingyun, et al.
Veröffentlicht: (2025)
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication Optimization
von: Tan, Zhanhong, et al.
Veröffentlicht: (2024)
von: Tan, Zhanhong, et al.
Veröffentlicht: (2024)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
GEM3D CIM General Purpose Matrix Computation Using 3D Integrated SRAM eDRAM Hybrid Compute In Memory on Memory Architecture
von: Chakraborty, Subhradip, et al.
Veröffentlicht: (2026)
von: Chakraborty, Subhradip, et al.
Veröffentlicht: (2026)
A Hybrid-Domain Floating-Point Compute-in-Memory Architecture for Efficient Acceleration of High-Precision Deep Neural Networks
von: Yi, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Yi, Zhiqiang, et al.
Veröffentlicht: (2025)
VeRA+: Vector-Based Lightweight Digital Compensation for Drift-Resilient RRAM In-Memory Computing
von: Dong, Weirong, et al.
Veröffentlicht: (2026)
von: Dong, Weirong, et al.
Veröffentlicht: (2026)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
Adaptive Hybrid FFT: A Novel Pipeline and Memory-Based Architecture for Radix-$2^k$ FFT in Large Size Processing
von: Zhao, Fangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Fangyu, et al.
Veröffentlicht: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
Limited Read-Write/Set Hardware Transactional Memory without modifying the ISA or the Coherence Protocol
von: Kafousis, Konstantinos
Veröffentlicht: (2025)
von: Kafousis, Konstantinos
Veröffentlicht: (2025)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
PULSE: Parametric Hardware Units for Low-power Sparsity-Aware Convolution Engine
von: Aliyev, Ilkin, et al.
Veröffentlicht: (2024)
von: Aliyev, Ilkin, et al.
Veröffentlicht: (2024)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
Efficient Page Migration in Hybrid Memory Systems
von: Upasna, et al.
Veröffentlicht: (2026)
von: Upasna, et al.
Veröffentlicht: (2026)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
Bi-SamplerZ: A Hardware-Efficient Gaussian Sampler Architecture for Quantum-Resistant Falcon Signatures
von: Zhao, Binke, et al.
Veröffentlicht: (2025)
von: Zhao, Binke, et al.
Veröffentlicht: (2025)
A Fully Pipelined FIFO Based Polynomial Multiplication Hardware Architecture Based On Number Theoretic Transform
von: Heidarpur, Moslem, et al.
Veröffentlicht: (2025)
von: Heidarpur, Moslem, et al.
Veröffentlicht: (2025)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
von: Shao, Haikuo, et al.
Veröffentlicht: (2024)
von: Shao, Haikuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025) -
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025) -
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025) -
A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025) -
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)