QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhixiong, Li, Haomin, Liu, Fangxin, Lu, Yuncheng, Wang, Zongwu, Yang, Tao, Jiang, Li, Guan, Haibing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LaMoS: Enabling Efficient Large Number Modular Multiplication through SRAM-based CiM Acceleration
by: Li, Haomin, et al.
Published: (2025)
by: Li, Haomin, et al.
Published: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
ASDR: Exploiting Adaptive Sampling and Data Reuse for CIM-based Instant Neural Rendering
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
ALLMod: Exploring $\underline{\mathbf{A}}$rea-Efficiency of $\underline{\mathbf{L}}$UT-based $\underline{\mathbf{L}}$arge Number $\underline{\mathbf{Mod}}$ular Reduction via Hybrid Workloads
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization
by: Wang, Zongwu, et al.
Published: (2025)
by: Wang, Zongwu, et al.
Published: (2025)
LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication
by: Wang, Zongwu, et al.
Published: (2024)
by: Wang, Zongwu, et al.
Published: (2024)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage Format
by: Zhao, Yilong, et al.
Published: (2025)
by: Zhao, Yilong, et al.
Published: (2025)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
by: Zhou, Xuwen, et al.
Published: (2026)
by: Zhou, Xuwen, et al.
Published: (2026)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
by: Liu, Fangxin, et al.
Published: (2026)
by: Liu, Fangxin, et al.
Published: (2026)
QUARK NOVA SIGNATURES IN SUPER-LUMINOUS SUPERNOVAE
by: M. Kostka
Published: (2014)
by: M. Kostka
Published: (2014)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
Accelerating Maximum Common Subgraph Computation by Exploiting Symmetries
by: Kothalawala, Buddhi, et al.
Published: (2026)
by: Kothalawala, Buddhi, et al.
Published: (2026)
Rectifiable Conductive Thermal Diodes Enabling Thermal Circuits with Selectable Operations for Thermal Logic Applications
by: Tian Li, et al.
Published: (2024)
by: Tian Li, et al.
Published: (2024)
Application-Oriented Benchmarking of Quantum Generative Learning Using QUARK
by: Kiwit, Florian J., et al.
Published: (2023)
by: Kiwit, Florian J., et al.
Published: (2023)
Long Noncoding RNA AF131217.1 Regulated Coronary Slow Flow-Induced Inflammation Affecting Coronary Slow Flow via KLF4
by: Haibing Jiang
Published: (2022)
by: Haibing Jiang
Published: (2022)
Enhancing Microwave-Optical Bell Pairs Generation for Quantum Transduction Using Kerr Nonlinearity
by: Li, Fangxin, et al.
Published: (2025)
by: Li, Fangxin, et al.
Published: (2025)
Accelerating Sparse Transformer Inference on GPU
by: Dai, Wenhao, et al.
Published: (2025)
by: Dai, Wenhao, et al.
Published: (2025)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
Fractional Operators for Nonlinear Electrical Circuits
by: Dassios, Ioannis
Published: (2025)
by: Dassios, Ioannis
Published: (2025)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
by: Jiang, Yanfeng, et al.
Published: (2024)
by: Jiang, Yanfeng, et al.
Published: (2024)
Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
by: Hu, Xiantao, et al.
Published: (2024)
by: Hu, Xiantao, et al.
Published: (2024)
QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers
by: Baras, Amit, et al.
Published: (2023)
by: Baras, Amit, et al.
Published: (2023)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
by: Wu, Chenpeng, et al.
Published: (2025)
by: Wu, Chenpeng, et al.
Published: (2025)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
by: Chen, Peilin, et al.
Published: (2025)
by: Chen, Peilin, et al.
Published: (2025)
Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination Enhancement
by: Kong, Derong, et al.
Published: (2025)
by: Kong, Derong, et al.
Published: (2025)
Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
by: Maisonnave, Lucas, et al.
Published: (2025)
by: Maisonnave, Lucas, et al.
Published: (2025)
Common Circuits
by: Murillo, Luis Felipe R.
Published: (2026)
by: Murillo, Luis Felipe R.
Published: (2026)
MORSE: An Efficient Homomorphic Secret Sharing Scheme Enabling Non-Linear Operation
by: Deng, Weiquan, et al.
Published: (2024)
by: Deng, Weiquan, et al.
Published: (2024)
Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
Benchmarking Quantum Generative Learning: A Study on Scalability and Noise Resilience using QUARK
by: Kiwit, Florian J., et al.
Published: (2024)
by: Kiwit, Florian J., et al.
Published: (2024)
QUARK: Robust Retrieval under Non-Faithful Queries via Query-Anchored Aggregation
by: Lyu, Rita Qiuran, et al.
Published: (2026)
by: Lyu, Rita Qiuran, et al.
Published: (2026)
Inductive Power Grid Cascading Failure Analysis with GRU-Gated Graph Attention
by: Zhou, Tianxin, et al.
Published: (2026)
by: Zhou, Tianxin, et al.
Published: (2026)
A Dual Power Grid Cascading Failure Model for the Vulnerability Analysis
by: Zhou, Tianxin, et al.
Published: (2024)
by: Zhou, Tianxin, et al.
Published: (2024)
FedCod: An Efficient Communication Protocol for Cross-Silo Federated Learning with Coding
by: Yan, Peishen, et al.
Published: (2024)
by: Yan, Peishen, et al.
Published: (2024)
Nebula: Enable City-Scale 3D Gaussian Splatting in Virtual Reality via Collaborative Rendering and Accelerated Stereo Rasterization
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Similar Items
-
LaMoS: Enabling Efficient Large Number Modular Multiplication through SRAM-based CiM Acceleration
by: Li, Haomin, et al.
Published: (2025) -
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025) -
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026) -
ASDR: Exploiting Adaptive Sampling and Data Reuse for CIM-based Instant Neural Rendering
by: Liu, Fangxin, et al.
Published: (2025) -
ALLMod: Exploring $\underline{\mathbf{A}}$rea-Efficiency of $\underline{\mathbf{L}}$UT-based $\underline{\mathbf{L}}$arge Number $\underline{\mathbf{Mod}}$ular Reduction via Hybrid Workloads
by: Liu, Fangxin, et al.
Published: (2025)