SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Parvathy, Aradhana Mohan, Ghosh, Soumendu Kumar, Kundu, Shamik, Raha, Arnab, Kundu, Souvik, Mathaikutty, Deepak A., Raghunathan, Anand |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
by: Raha, Arnab, et al.
Published: (2024)
by: Raha, Arnab, et al.
Published: (2024)
GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
by: Das, Arghadip, et al.
Published: (2025)
by: Das, Arghadip, et al.
Published: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025)
by: Wu, Michael, et al.
Published: (2025)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
by: Bhattacharya, Swastik, et al.
Published: (2025)
by: Bhattacharya, Swastik, et al.
Published: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
XAMBA: Enabling Efficient State Space Models on Resource-Constrained Neural Processing Units
by: Das, Arghadip, et al.
Published: (2025)
by: Das, Arghadip, et al.
Published: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
TYTAN: Taylor-series based Non-Linear Activation Engine for Deep Learning Accelerators
by: Pramanik, Soham, et al.
Published: (2025)
by: Pramanik, Soham, et al.
Published: (2025)
SiTe CiM: Signed Ternary Computing-in-Memory for Ultra-Low Precision Deep Neural Networks
by: Thakuria, Niharika, et al.
Published: (2024)
by: Thakuria, Niharika, et al.
Published: (2024)
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
by: Kundu, Souvik, et al.
Published: (2024)
by: Kundu, Souvik, et al.
Published: (2024)
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
by: Lin, Zi-Wei, et al.
Published: (2026)
by: Lin, Zi-Wei, et al.
Published: (2026)
Technology solutions targeting the performance of gen-AI inference in resource constrained platforms
by: Kundu, Joyjit, et al.
Published: (2026)
by: Kundu, Joyjit, et al.
Published: (2026)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
by: Yu, Feng, et al.
Published: (2026)
by: Yu, Feng, et al.
Published: (2026)
A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence
by: Kundu, Yudhishthira, et al.
Published: (2025)
by: Kundu, Yudhishthira, et al.
Published: (2025)
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
by: Zheng, Keran, et al.
Published: (2025)
by: Zheng, Keran, et al.
Published: (2025)
KLLM: Fast LLM Inference with K-Means Quantization
by: Wu, Xueying, et al.
Published: (2025)
by: Wu, Xueying, et al.
Published: (2025)
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization
by: Kim, Daeun, et al.
Published: (2025)
by: Kim, Daeun, et al.
Published: (2025)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
In-Memory ADC-Based Nonlinear Activation Quantization for Efficient In-Memory Computing
by: Dong, Shuai, et al.
Published: (2026)
by: Dong, Shuai, et al.
Published: (2026)
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
by: Kim, Seeyeon, et al.
Published: (2026)
by: Kim, Seeyeon, et al.
Published: (2026)
Memory Scraping Attack on Xilinx FPGAs: Private Data Extraction from Terminated Processes
by: Madabhushi, Bharadwaj, et al.
Published: (2024)
by: Madabhushi, Bharadwaj, et al.
Published: (2024)
Pushing the Limits of BFP on Narrow Precision LLM Inference
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
by: Vungarala, Deepak, et al.
Published: (2025)
by: Vungarala, Deepak, et al.
Published: (2025)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
by: Vishwanathan, Manoj, et al.
Published: (2026)
by: Vishwanathan, Manoj, et al.
Published: (2026)
System-performance and cost modeling of Large Language Model training and inference
by: Guo, Wenzhe, et al.
Published: (2025)
by: Guo, Wenzhe, et al.
Published: (2025)
A3D: Agentic AI flow for autonomous Accelerator Design
by: Nallathambi, Abinand, et al.
Published: (2026)
by: Nallathambi, Abinand, et al.
Published: (2026)
Guardians of the Quantum GAN
by: Ghosh, Archisman, et al.
Published: (2024)
by: Ghosh, Archisman, et al.
Published: (2024)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
by: Chen, Peilin, et al.
Published: (2025)
by: Chen, Peilin, et al.
Published: (2025)
Towards Efficient Design Verification -- Constrained Random Verification using PyUVM
by: Gadde, Deepak Narayan, et al.
Published: (2024)
by: Gadde, Deepak Narayan, et al.
Published: (2024)
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference
by: Chen, Guoci, et al.
Published: (2026)
by: Chen, Guoci, et al.
Published: (2026)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
SPEED: A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
by: Yu, Zhongkai, et al.
Published: (2024)
by: Yu, Zhongkai, et al.
Published: (2024)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
by: Yin, Chenyang, et al.
Published: (2025)
by: Yin, Chenyang, et al.
Published: (2025)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
by: Su, Zhongling, et al.
Published: (2025)
by: Su, Zhongling, et al.
Published: (2025)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
by: Zheng, Jianing, et al.
Published: (2025)
by: Zheng, Jianing, et al.
Published: (2025)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
by: Heo, Guseul, et al.
Published: (2024)
by: Heo, Guseul, et al.
Published: (2024)
Similar Items
-
FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
by: Raha, Arnab, et al.
Published: (2024) -
GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
by: Das, Arghadip, et al.
Published: (2025) -
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025) -
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025) -
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
by: Bhattacharya, Swastik, et al.
Published: (2025)