31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Pingcheng, Tan, Yonghao, Liu, Xuejiao, Luo, Peng, Liu, Yu, Pang, Di, Ma, Songchen, Huang, Xijie, Liu, Shih-Yang, Zhang, Dong, Lu, Zhichao, Liang, Luhong, Tsui, Chi-Ying, Tu, Fengbin, Zhao, Liang, Cheng, Kwang-Ting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
von: Tan, Yonghao, et al.
Veröffentlicht: (2025)
von: Tan, Yonghao, et al.
Veröffentlicht: (2025)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
A 28nm 0.22μJ/token memory-compute-intensity-aware CNN-Transformer accelerator with hybrid-attention-based layer-fusion and cascaded pruning for semantic-segmentation
von: Dong, Pingcheng, et al.
Veröffentlicht: (2025)
von: Dong, Pingcheng, et al.
Veröffentlicht: (2025)
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
von: Dong, Pingcheng, et al.
Veröffentlicht: (2024)
von: Dong, Pingcheng, et al.
Veröffentlicht: (2024)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
von: Shao, Kunming, et al.
Veröffentlicht: (2025)
von: Shao, Kunming, et al.
Veröffentlicht: (2025)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
Ising-ReRAM: A Low Power Ising Machine ReRAM Crossbar for NP Problems
von: Bloomer, Everest, et al.
Veröffentlicht: (2026)
von: Bloomer, Everest, et al.
Veröffentlicht: (2026)
All-in-Memory Stochastic Computing using ReRAM
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2025)
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2025)
Efficient and Robust Quantization-aware Training via Adaptive Coreset Selection
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
Memory Efficient Transformer Adapter for Dense Predictions
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
Towards Customized Knowledge Distillation for Chip-Level Dense Image Predictions
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
DFlash: Block Diffusion for Flash Speculative Decoding
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
A Fully Automated Platform for Evaluating ReRAM Crossbars
von: Pelke, Rebecca, et al.
Veröffentlicht: (2024)
von: Pelke, Rebecca, et al.
Veröffentlicht: (2024)
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
von: Li, Bingbing, et al.
Veröffentlicht: (2024)
von: Li, Bingbing, et al.
Veröffentlicht: (2024)
Device Modeling Bias in ReRAM-based Neural Network Simulations
von: Yousuf, Osama, et al.
Veröffentlicht: (2022)
von: Yousuf, Osama, et al.
Veröffentlicht: (2022)
ReRAM/CMOS Array Integration and Characterization via Design of Experiments
von: Imtiaz Hossen, et al.
Veröffentlicht: (2025)
von: Imtiaz Hossen, et al.
Veröffentlicht: (2025)
ARAS: An Adaptive Low-Cost ReRAM-Based Accelerator for DNNs
von: Sabri, Mohammad, et al.
Veröffentlicht: (2024)
von: Sabri, Mohammad, et al.
Veröffentlicht: (2024)
Characterizing the range of the complex Monge-Ampère operator
von: Liu, Songchen
Veröffentlicht: (2024)
von: Liu, Songchen
Veröffentlicht: (2024)
Sensitivity-Aware Mixed-Precision Quantization for ReRAM-based Computing-in-Memory
von: Chen, Guan-Cheng, et al.
Veröffentlicht: (2025)
von: Chen, Guan-Cheng, et al.
Veröffentlicht: (2025)
Hiding Information for Secure and Covert Data Storage in Commercial ReRAM Chips
von: Ferdaus, Farah, et al.
Veröffentlicht: (2024)
von: Ferdaus, Farah, et al.
Veröffentlicht: (2024)
HURRY: Highly Utilized, Reconfigurable ReRAM-based In-situ Accelerator with Multifunctionality
von: Shin, Hery, et al.
Veröffentlicht: (2024)
von: Shin, Hery, et al.
Veröffentlicht: (2024)
Online Soft Error Tolerance in ReRAM Crossbars for Deep Learning Accelerators
von: Khezeli, Benyamin, et al.
Veröffentlicht: (2024)
von: Khezeli, Benyamin, et al.
Veröffentlicht: (2024)
FARe: Fault-Aware GNN Training on ReRAM-based PIM Accelerators
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2024)
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2024)
Edge Training and Inference with Analog ReRAM Technology for Hand Gesture Recognition
von: Clerico, Victoria, et al.
Veröffentlicht: (2025)
von: Clerico, Victoria, et al.
Veröffentlicht: (2025)
Hamun: An Approximate Computation Method to Prolong the Lifespan of ReRAM-Based Accelerators
von: Sabri, Mohammad, et al.
Veröffentlicht: (2025)
von: Sabri, Mohammad, et al.
Veröffentlicht: (2025)
Design Space Exploration for ReRAM-based Architectures to Address Scaling Non-idealities
von: Lin, Ching-Yi, et al.
Veröffentlicht: (2026)
von: Lin, Ching-Yi, et al.
Veröffentlicht: (2026)
Conductive metal oxide and hafnium oxide bilayer ReRAM: an ab initio study
von: Honet, Antoine, et al.
Veröffentlicht: (2024)
von: Honet, Antoine, et al.
Veröffentlicht: (2024)
Stuck-at Faults in ReRAM Neuromorphic Circuit Array and their Correction through Machine Learning
von: Sawal, Vedant, et al.
Veröffentlicht: (2024)
von: Sawal, Vedant, et al.
Veröffentlicht: (2024)
PdNeuRAM: forming-free, multi-bit Pd/HfO2 ReRAM for energy-efficient neuromorphic computing
von: Hua, Erbing, et al.
Veröffentlicht: (2025)
von: Hua, Erbing, et al.
Veröffentlicht: (2025)
Multi-Objective Optimization of ReRAM Crossbars for Robust DNN Inferencing under Stochastic Noise
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2021)
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2021)
A Fully Hardware Implemented Accelerator Design in ReRAM Analog Computing without ADCs
von: Dang, Peng, et al.
Veröffentlicht: (2024)
von: Dang, Peng, et al.
Veröffentlicht: (2024)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
von: Jiang, Xixi, et al.
Veröffentlicht: (2025)
von: Jiang, Xixi, et al.
Veröffentlicht: (2025)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
von: Zhang, Chenguang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenguang, et al.
Veröffentlicht: (2024)
ReCross: Efficient Embedding Reduction Scheme for In-Memory Computing using ReRAM-Based Crossbar
von: Lai, Yu-Hong, et al.
Veröffentlicht: (2025)
von: Lai, Yu-Hong, et al.
Veröffentlicht: (2025)
Analytical Modelling of the Transport in Analog Filamentary Conductive-Metal-Oxide/HfOx ReRAM Devices
von: Falcone, Donato Francesco, et al.
Veröffentlicht: (2025)
von: Falcone, Donato Francesco, et al.
Veröffentlicht: (2025)
Hardware Implementation of Ring Oscillator Networks Coupled by BEOL Integrated ReRAM for Associative Memory Tasks
von: Choi, Wooseok, et al.
Veröffentlicht: (2025)
von: Choi, Wooseok, et al.
Veröffentlicht: (2025)
SR-LIO++: Efficient LiDAR-Inertial Odometry and Quantized Mapping with Sweep Reconstruction
von: Yuan, Zikang, et al.
Veröffentlicht: (2025)
von: Yuan, Zikang, et al.
Veröffentlicht: (2025)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
von: Tan, Yonghao, et al.
Veröffentlicht: (2025) -
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
von: Ma, Songchen, et al.
Veröffentlicht: (2026) -
A 28nm 0.22μJ/token memory-compute-intensity-aware CNN-Transformer accelerator with hybrid-attention-based layer-fusion and cascaded pruning for semantic-segmentation
von: Dong, Pingcheng, et al.
Veröffentlicht: (2025) -
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
von: Dong, Pingcheng, et al.
Veröffentlicht: (2024) -
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)