JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Mingzi, Meng, Yuan, Tang, Chen, Zhang, Weixiang, Qin, Yijian, Yao, Yang, Li, Yingxin, Feng, Tongtong, Wang, Xin, Guan, Xun, Wang, Zhi, Zhu, Wenwu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
di: Xie, Xilong, et al.
Pubblicazione: (2025)
di: Xie, Xilong, et al.
Pubblicazione: (2025)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
di: Mo, Zhiwen, et al.
Pubblicazione: (2024)
di: Mo, Zhiwen, et al.
Pubblicazione: (2024)
NAS-Bench-Graph: Benchmarking Graph Neural Architecture Search
di: Qin, Yijian, et al.
Pubblicazione: (2022)
di: Qin, Yijian, et al.
Pubblicazione: (2022)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
Multi-weather Cross-view Geo-localization Using Denoising Diffusion Models
di: Feng, Tongtong, et al.
Pubblicazione: (2024)
di: Feng, Tongtong, et al.
Pubblicazione: (2024)
TMPQ-DM: Joint Timestep Reduction and Quantization Precision Selection for Efficient Diffusion Models
di: Sun, Haojun, et al.
Pubblicazione: (2024)
di: Sun, Haojun, et al.
Pubblicazione: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
Disentangled Representation Learning with Large Language Models for Text-Attributed Graphs
di: Qin, Yijian, et al.
Pubblicazione: (2023)
di: Qin, Yijian, et al.
Pubblicazione: (2023)
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantization
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
Enhancing Implicit Neural Representations via Symmetric Power Transformation
di: Zhang, Weixiang, et al.
Pubblicazione: (2024)
di: Zhang, Weixiang, et al.
Pubblicazione: (2024)
Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective
di: Jiang, Jiacheng, et al.
Pubblicazione: (2025)
di: Jiang, Jiacheng, et al.
Pubblicazione: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
di: You, Dean, et al.
Pubblicazione: (2025)
di: You, Dean, et al.
Pubblicazione: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
Causal-aware Graph Neural Architecture Search under Distribution Shifts
di: Li, Peiwen, et al.
Pubblicazione: (2024)
di: Li, Peiwen, et al.
Pubblicazione: (2024)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
di: Chen, Lei, et al.
Pubblicazione: (2024)
di: Chen, Lei, et al.
Pubblicazione: (2024)
EVOS: Efficient Implicit Neural Training via EVOlutionary Selector
di: Zhang, Weixiang, et al.
Pubblicazione: (2024)
di: Zhang, Weixiang, et al.
Pubblicazione: (2024)
Self-evolving Embodied AI
di: Feng, Tongtong, et al.
Pubblicazione: (2026)
di: Feng, Tongtong, et al.
Pubblicazione: (2026)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
di: Wang, YiFeng, et al.
Pubblicazione: (2026)
di: Wang, YiFeng, et al.
Pubblicazione: (2026)
LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design
di: Wei, Renjie, et al.
Pubblicazione: (2025)
di: Wei, Renjie, et al.
Pubblicazione: (2025)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
di: Wang, Zhao, et al.
Pubblicazione: (2021)
di: Wang, Zhao, et al.
Pubblicazione: (2021)
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
di: Feng, Weilun, et al.
Pubblicazione: (2024)
di: Feng, Weilun, et al.
Pubblicazione: (2024)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
di: Su, Le, et al.
Pubblicazione: (2026)
di: Su, Le, et al.
Pubblicazione: (2026)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
di: Zhang, Xi, et al.
Pubblicazione: (2025)
di: Zhang, Xi, et al.
Pubblicazione: (2025)
The Case for Co-Designing Model Architectures with Hardware
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
di: Li, Zhen, et al.
Pubblicazione: (2025)
di: Li, Zhen, et al.
Pubblicazione: (2025)
Post-training Quantization for Text-to-Image Diffusion Models with Progressive Calibration and Activation Relaxing
di: Tang, Siao, et al.
Pubblicazione: (2023)
di: Tang, Siao, et al.
Pubblicazione: (2023)
Retraining-free Model Quantization via One-Shot Weight-Coupling Learning
di: Tang, Chen, et al.
Pubblicazione: (2024)
di: Tang, Chen, et al.
Pubblicazione: (2024)
Behavior Importance-Aware Graph Neural Architecture Search for Cross-Domain Recommendation
di: Ge, Chendi, et al.
Pubblicazione: (2025)
di: Ge, Chendi, et al.
Pubblicazione: (2025)
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
di: Zhang, Yuanpeng, et al.
Pubblicazione: (2025)
di: Zhang, Yuanpeng, et al.
Pubblicazione: (2025)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
Low-Bit Quantization of Bandlimited Graph Signals via Iterative Methods
di: Krahmer, Felix, et al.
Pubblicazione: (2026)
di: Krahmer, Felix, et al.
Pubblicazione: (2026)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
di: Makin, Yashasvi, et al.
Pubblicazione: (2025)
di: Makin, Yashasvi, et al.
Pubblicazione: (2025)
LLM4DyG: Can Large Language Models Solve Spatial-Temporal Problems on Dynamic Graphs?
di: Zhang, Zeyang, et al.
Pubblicazione: (2023)
di: Zhang, Zeyang, et al.
Pubblicazione: (2023)
Hardware-Software Co-Design for Event-Driven SNN Deployment on Low-Cost Neuromorphic FPGAs
di: Lee, Jiwoon, et al.
Pubblicazione: (2026)
di: Lee, Jiwoon, et al.
Pubblicazione: (2026)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
di: Lee, Banseok, et al.
Pubblicazione: (2025)
di: Lee, Banseok, et al.
Pubblicazione: (2025)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Towards Secure and Efficient DNN Accelerators via Hardware-Software Co-Design
di: Xuan, Wei, et al.
Pubblicazione: (2026)
di: Xuan, Wei, et al.
Pubblicazione: (2026)
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
di: Deng, Kaiyuan, et al.
Pubblicazione: (2026)
di: Deng, Kaiyuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
di: Xie, Xilong, et al.
Pubblicazione: (2025) -
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
di: Mo, Zhiwen, et al.
Pubblicazione: (2024) -
NAS-Bench-Graph: Benchmarking Graph Neural Architecture Search
di: Qin, Yijian, et al.
Pubblicazione: (2022) -
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
di: Zhang, Yu, et al.
Pubblicazione: (2024) -
Multi-weather Cross-view Geo-localization Using Denoising Diffusion Models
di: Feng, Tongtong, et al.
Pubblicazione: (2024)