FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Xilong, Wang, Liang, Xiao, Limin, Han, Meng, Sun, Lin, Zheng, Shuai, Xu, Xiangrong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
by: Han, Meng, et al.
Published: (2023)
by: Han, Meng, et al.
Published: (2023)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
by: Han, Meng, et al.
Published: (2023)
by: Han, Meng, et al.
Published: (2023)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
by: Dai, Xilai, et al.
Published: (2024)
by: Dai, Xilai, et al.
Published: (2024)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
by: Lu, Tsung-Han, et al.
Published: (2026)
by: Lu, Tsung-Han, et al.
Published: (2026)
Hardware-Software Co-Design for Event-Driven SNN Deployment on Low-Cost Neuromorphic FPGAs
by: Lee, Jiwoon, et al.
Published: (2026)
by: Lee, Jiwoon, et al.
Published: (2026)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
by: Lu, Tsung-Han, et al.
Published: (2025)
by: Lu, Tsung-Han, et al.
Published: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)
by: You, Dean, et al.
Published: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
C2HLSC: Can LLMs Bridge the Software-to-Hardware Design Gap?
by: Collini, Luca, et al.
Published: (2024)
by: Collini, Luca, et al.
Published: (2024)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
by: Wang, Zhao, et al.
Published: (2021)
by: Wang, Zhao, et al.
Published: (2021)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
by: Kim, Dong Eun, et al.
Published: (2025)
by: Kim, Dong Eun, et al.
Published: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)
by: Zhou, Zhe, et al.
Published: (2024)
Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-Design
by: Pan, Tianwei, et al.
Published: (2025)
by: Pan, Tianwei, et al.
Published: (2025)
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025)
by: Tan, Yonghao, et al.
Published: (2025)
Aquas: Enhancing Domain Specialization through Holistic Hardware-Software Co-Optimization based on MLIR
by: Zou, Yuyang, et al.
Published: (2025)
by: Zou, Yuyang, et al.
Published: (2025)
SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
by: Wang, Chuanzhen, et al.
Published: (2026)
by: Wang, Chuanzhen, et al.
Published: (2026)
Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
by: Kanda, R., et al.
Published: (2026)
by: Kanda, R., et al.
Published: (2026)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
by: Li, Cong, et al.
Published: (2026)
by: Li, Cong, et al.
Published: (2026)
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025)
by: Wu, Michael, et al.
Published: (2025)
Hardware-Aware Fine-Tuning of Spiking Q-Networks on the SpiNNaker2 Neuromorphic Platform
by: Arfa, Sirine, et al.
Published: (2025)
by: Arfa, Sirine, et al.
Published: (2025)
Annotating Slack Directly on Your Verilog: Fine-Grained RTL Timing Evaluation for Early Optimization
by: Fang, Wenji, et al.
Published: (2024)
by: Fang, Wenji, et al.
Published: (2024)
A Prototype-Based Framework to Design Scalable Heterogeneous SoCs with Fine-Grained DFS
by: Montanaro, Gabriele, et al.
Published: (2024)
by: Montanaro, Gabriele, et al.
Published: (2024)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
Design Environment of Quantization-Aware Edge AI Hardware for Few-Shot Learning
by: Kanda, R., et al.
Published: (2026)
by: Kanda, R., et al.
Published: (2026)
TTP: A Hardware-Efficient Design for Precise Prefetching in Ray Tracing
by: Tozlu, Yavuz Selim, et al.
Published: (2026)
by: Tozlu, Yavuz Selim, et al.
Published: (2026)
AssertLLM: Generating Hardware Verification Assertions from Design Specifications via Multi-LLMs
by: Yan, Zhiyuan, et al.
Published: (2024)
by: Yan, Zhiyuan, et al.
Published: (2024)
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
by: Fyon, Arthur, et al.
Published: (2026)
by: Fyon, Arthur, et al.
Published: (2026)
Sparsity-Aware Hardware-Software Co-Design of Spiking Neural Networks: An Overview
by: Aliyev, Ilkin, et al.
Published: (2024)
by: Aliyev, Ilkin, et al.
Published: (2024)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
by: Raj, Ritik, et al.
Published: (2025)
by: Raj, Ritik, et al.
Published: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
by: Alsop, Johnathan, et al.
Published: (2023)
by: Alsop, Johnathan, et al.
Published: (2023)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
AssertLLM: Generating and Evaluating Hardware Verification Assertions from Design Specifications via Multi-LLMs
by: Fang, Wenji, et al.
Published: (2024)
by: Fang, Wenji, et al.
Published: (2024)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
by: Zhou, Zikai, et al.
Published: (2025)
by: Zhou, Zikai, et al.
Published: (2025)
Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
by: Liao, Yuchao, et al.
Published: (2026)
by: Liao, Yuchao, et al.
Published: (2026)
DreamRAM: A Fine-Grained Configurable Design Space Modeling Tool for Custom 3D Die-Stacked DRAM
by: Cai, Victor, et al.
Published: (2025)
by: Cai, Victor, et al.
Published: (2025)
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
by: Shin, Changmin, et al.
Published: (2025)
by: Shin, Changmin, et al.
Published: (2025)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
by: Nguyen, Hoa, et al.
Published: (2025)
by: Nguyen, Hoa, et al.
Published: (2025)
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
by: Tsai, Pei-Huan, et al.
Published: (2026)
by: Tsai, Pei-Huan, et al.
Published: (2026)
Similar Items
-
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
by: Han, Meng, et al.
Published: (2023) -
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025) -
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
by: Han, Meng, et al.
Published: (2023) -
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
by: Dai, Xilai, et al.
Published: (2024) -
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
by: Lu, Tsung-Han, et al.
Published: (2026)