AHCQ-SAM: Toward Accurate and Hardware-Compatible Post-Training Segment Anything Model Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Wenlun, Zhong, Yunshan, Yan, Weiqi, Zhang, Shengchuan, Ando, Shimpei, Yoshioka, Kentaro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM Circuits
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024)
A Review of SRAM-based Compute-in-Memory Circuits
von: Yoshioka, Kentaro, et al.
Veröffentlicht: (2024)
von: Yoshioka, Kentaro, et al.
Veröffentlicht: (2024)
PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
von: Kida, Misaki, et al.
Veröffentlicht: (2025)
von: Kida, Misaki, et al.
Veröffentlicht: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
von: Killian, Earl
Veröffentlicht: (2026)
von: Killian, Earl
Veröffentlicht: (2026)
HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
von: Zhang, Yipu, et al.
Veröffentlicht: (2025)
von: Zhang, Yipu, et al.
Veröffentlicht: (2025)
Energy-Efficient Hardware Acceleration of Whisper ASR on a CGLA
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
FireBridge: Cycle-Accurate Hardware + Firmware Co-Verification for Modern Accelerators
von: Abarajithan, G, et al.
Veröffentlicht: (2026)
von: Abarajithan, G, et al.
Veröffentlicht: (2026)
Design Environment of Quantization-Aware Edge AI Hardware for Few-Shot Learning
von: Kanda, R., et al.
Veröffentlicht: (2026)
von: Kanda, R., et al.
Veröffentlicht: (2026)
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
von: Li, Shuaiting, et al.
Veröffentlicht: (2024)
von: Li, Shuaiting, et al.
Veröffentlicht: (2024)
Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
von: Popovici, Doru Thom, et al.
Veröffentlicht: (2025)
von: Popovici, Doru Thom, et al.
Veröffentlicht: (2025)
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
von: Tan, Yonghao, et al.
Veröffentlicht: (2025)
von: Tan, Yonghao, et al.
Veröffentlicht: (2025)
AutoPDR: Circuit-Aware Solver Configuration Prediction for Hardware Model Checking
von: Hu, Guangyu, et al.
Veröffentlicht: (2026)
von: Hu, Guangyu, et al.
Veröffentlicht: (2026)
Hardware-Compatible Single-Shot Feasible-Space Heuristics for Solving the Quadratic Assignment Problem
von: Im, Haesol, et al.
Veröffentlicht: (2025)
von: Im, Haesol, et al.
Veröffentlicht: (2025)
Improving Quantization with Post-Training Model Expansion
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
Towards the Certification of Hybrid Architectures: Analysing Interference on Hardware Accelerators through PML
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
AssertLLM: Generating Hardware Verification Assertions from Design Specifications via Multi-LLMs
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2024)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
von: Zhang, Kunlong, et al.
Veröffentlicht: (2025)
von: Zhang, Kunlong, et al.
Veröffentlicht: (2025)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
AssertLLM: Generating and Evaluating Hardware Verification Assertions from Design Specifications via Multi-LLMs
von: Fang, Wenji, et al.
Veröffentlicht: (2024)
von: Fang, Wenji, et al.
Veröffentlicht: (2024)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
Facial Expression Recognition System Using DNN Accelerator with Multi-threading on FPGA
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification
von: Zhong, Yang, et al.
Veröffentlicht: (2025)
von: Zhong, Yang, et al.
Veröffentlicht: (2025)
Gen-NeRF: Efficient and Generalizable Neural Radiance Fields via Algorithm-Hardware Co-Design
von: Fu, Yonggan, et al.
Veröffentlicht: (2023)
von: Fu, Yonggan, et al.
Veröffentlicht: (2023)
Efficient yet Accurate End-to-End SC Accelerator Design
von: Li, Meng, et al.
Veröffentlicht: (2024)
von: Li, Meng, et al.
Veröffentlicht: (2024)
Invited: Toward Accurate, Large-scale Electromigration Analysis and Optimization in Integrated Systems
von: Sapatnekar, Sachin S.
Veröffentlicht: (2026)
von: Sapatnekar, Sachin S.
Veröffentlicht: (2026)
Towards Closing the Performance Gap for Cryptographic Kernels Between CPUs and Specialized Hardware
von: Zhang, Naifeng, et al.
Veröffentlicht: (2025)
von: Zhang, Naifeng, et al.
Veröffentlicht: (2025)
Towards Efficient and Accurate Detection of On-Chip Fail-Slow Failures for Many-Core Accelerators
von: Wu, Junchi, et al.
Veröffentlicht: (2025)
von: Wu, Junchi, et al.
Veröffentlicht: (2025)
Sustainable Hardware Specialization
von: Dangi, Pranav, et al.
Veröffentlicht: (2024)
von: Dangi, Pranav, et al.
Veröffentlicht: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
Linear Complexity Fermionic Simulation on Quantum Devices with Hardware Connectivity Constraints
von: Gao, Xiangyu, et al.
Veröffentlicht: (2026)
von: Gao, Xiangyu, et al.
Veröffentlicht: (2026)
elasticAI.explorer: Towards a Unified End-to-End Framework for Hardware-Aware Neural Architecture Search
von: Maman, Natalie, et al.
Veröffentlicht: (2026)
von: Maman, Natalie, et al.
Veröffentlicht: (2026)
Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
von: Wu, Meng, et al.
Veröffentlicht: (2024)
von: Wu, Meng, et al.
Veröffentlicht: (2024)
Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
von: Carrigg, Kieran, et al.
Veröffentlicht: (2026)
von: Carrigg, Kieran, et al.
Veröffentlicht: (2026)
Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025) -
ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM Circuits
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024) -
A Review of SRAM-based Compute-in-Memory Circuits
von: Yoshioka, Kentaro, et al.
Veröffentlicht: (2024) -
PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024) -
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
von: Kida, Misaki, et al.
Veröffentlicht: (2025)