Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Gou, Terry, Gupta, Puneet |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
by: Rashidi, Saeed, et al.
Published: (2024)
by: Rashidi, Saeed, et al.
Published: (2024)
Efficient Implementation of LinearUCB through Algorithmic Improvements and Vector Computing Acceleration for Embedded Learning Systems
by: Angioli, Marco, et al.
Published: (2025)
by: Angioli, Marco, et al.
Published: (2025)
CIMPool: Scalable Neural Network Acceleration for Compute-In-Memory using Weight Pools
by: Li, Shurui, et al.
Published: (2025)
by: Li, Shurui, et al.
Published: (2025)
MOGNET: A Mux-residual quantized Network leveraging Online-Generated weights
by: Nguyen, Van Thien, et al.
Published: (2025)
by: Nguyen, Van Thien, et al.
Published: (2025)
ARTEMIS: A Mixed Analog-Stochastic In-DRAM Accelerator for Transformer Neural Networks
by: Afifi, Salma, et al.
Published: (2024)
by: Afifi, Salma, et al.
Published: (2024)
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
by: Sun, Ruiqi, et al.
Published: (2023)
by: Sun, Ruiqi, et al.
Published: (2023)
Efficient and Mathematically Robust Operations for Certified Neural Networks Inference
by: Geyer, Fabien, et al.
Published: (2024)
by: Geyer, Fabien, et al.
Published: (2024)
MiCo: End-to-End Mixed Precision Neural Network Co-Exploration Framework for Edge AI
by: Jiang, Zijun, et al.
Published: (2025)
by: Jiang, Zijun, et al.
Published: (2025)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
by: Fang, Jiaxun, et al.
Published: (2025)
by: Fang, Jiaxun, et al.
Published: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
Efficient Data Access Paths for Mixed Vector-Relational Search
by: Sanca, Viktor, et al.
Published: (2024)
by: Sanca, Viktor, et al.
Published: (2024)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
by: Gupta, Neelesh, et al.
Published: (2026)
by: Gupta, Neelesh, et al.
Published: (2026)
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
by: Rutishauser, Georg, et al.
Published: (2024)
by: Rutishauser, Georg, et al.
Published: (2024)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
by: Khataei, Alireza, et al.
Published: (2025)
by: Khataei, Alireza, et al.
Published: (2025)
Efficient and Reliable Vector Similarity Search Using Asymmetric Encoding with NAND-Flash for Many-Class Few-Shot Learning
by: Chiang, Hao-Wei, et al.
Published: (2024)
by: Chiang, Hao-Wei, et al.
Published: (2024)
An FPGA-Based SoC Architecture with a RISC-V Controller for Energy-Efficient Temporal-Coding Spiking Neural Networks
by: Sekonji, Mohammad Javad, et al.
Published: (2026)
by: Sekonji, Mohammad Javad, et al.
Published: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
ROSA: Robust and Energy-Efficient Microring-Based Optical Neural Networks via Optical Shift-and-Add and Layer-Wise Hybrid Mapping
by: Zhang, Huifan, et al.
Published: (2026)
by: Zhang, Huifan, et al.
Published: (2026)
Learning Library Cell Representations in Vector Space
by: Liang, Rongjian, et al.
Published: (2025)
by: Liang, Rongjian, et al.
Published: (2025)
Sustainable Transformer Neural Network Acceleration with Stochastic Photonic Computing
by: Afifi, S., et al.
Published: (2026)
by: Afifi, S., et al.
Published: (2026)
Approximate Multiplier Induced Error Propagation in Deep Neural Networks
by: Alahakoon, A. M. H. H., et al.
Published: (2025)
by: Alahakoon, A. M. H. H., et al.
Published: (2025)
Histogram-Equalized Quantization for logic-gated Residual Neural Networks
by: Nguyen, Van Thien, et al.
Published: (2025)
by: Nguyen, Van Thien, et al.
Published: (2025)
Graph Neural Networks Based Analog Circuit Link Prediction
by: Pan, Guanyuan, et al.
Published: (2025)
by: Pan, Guanyuan, et al.
Published: (2025)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
by: Wu, Ka Wai
Published: (2024)
by: Wu, Ka Wai
Published: (2024)
MCU-MixQ: A HW/SW Co-optimized Mixed-precision Neural Network Design Framework for MCUs
by: Gong, Junfeng, et al.
Published: (2024)
by: Gong, Junfeng, et al.
Published: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
by: Gimenes, Pedro, et al.
Published: (2025)
by: Gimenes, Pedro, et al.
Published: (2025)
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization
by: Kim, Daeun, et al.
Published: (2025)
by: Kim, Daeun, et al.
Published: (2025)
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
by: Alama, Youssef A. Ait, et al.
Published: (2024)
by: Alama, Youssef A. Ait, et al.
Published: (2024)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
by: Andronic, Marta, et al.
Published: (2024)
by: Andronic, Marta, et al.
Published: (2024)
PhotoGAN: Generative Adversarial Neural Network Acceleration with Silicon Photonics
by: Suresh, Tharini, et al.
Published: (2025)
by: Suresh, Tharini, et al.
Published: (2025)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
by: Yan, Minghao, et al.
Published: (2023)
by: Yan, Minghao, et al.
Published: (2023)
A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE
by: Okubo, Ikumi, et al.
Published: (2024)
by: Okubo, Ikumi, et al.
Published: (2024)
Accelerating Neural Networks for Large Language Models and Graph Processing with Silicon Photonics
by: Afifi, Salma, et al.
Published: (2024)
by: Afifi, Salma, et al.
Published: (2024)
FuncGNN: Learning Functional Semantics of Logic Circuits with Graph Neural Networks
by: Zhao, Qiyun
Published: (2025)
by: Zhao, Qiyun
Published: (2025)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
by: Klhufek, Jan, et al.
Published: (2024)
by: Klhufek, Jan, et al.
Published: (2024)
The prediction of the quality of results in Logic Synthesis using Transformer and Graph Neural Networks
by: Yang, Chenghao, et al.
Published: (2022)
by: Yang, Chenghao, et al.
Published: (2022)
MCEL: Margin-Based Cross-Entropy Loss for Error-Tolerant Quantized Neural Networks
by: Yayla, Mikail, et al.
Published: (2026)
by: Yayla, Mikail, et al.
Published: (2026)
Boolean-aware Boolean Circuit Classification: A Comprehensive Study on Graph Neural Network
by: Ni, Liwei, et al.
Published: (2024)
by: Ni, Liwei, et al.
Published: (2024)
Optical Computing for Deep Neural Network Acceleration: Foundations, Recent Developments, and Emerging Directions
by: Pasricha, Sudeep
Published: (2024)
by: Pasricha, Sudeep
Published: (2024)
AiEDA: An Open-Source AI-Aided Design Library for Design-to-Vector
by: Qiu, Yihang, et al.
Published: (2025)
by: Qiu, Yihang, et al.
Published: (2025)
Similar Items
-
FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
by: Rashidi, Saeed, et al.
Published: (2024) -
Efficient Implementation of LinearUCB through Algorithmic Improvements and Vector Computing Acceleration for Embedded Learning Systems
by: Angioli, Marco, et al.
Published: (2025) -
CIMPool: Scalable Neural Network Acceleration for Compute-In-Memory using Weight Pools
by: Li, Shurui, et al.
Published: (2025) -
MOGNET: A Mux-residual quantized Network leveraging Online-Generated weights
by: Nguyen, Van Thien, et al.
Published: (2025) -
ARTEMIS: A Mixed Analog-Stochastic In-DRAM Accelerator for Transformer Neural Networks
by: Afifi, Salma, et al.
Published: (2024)