Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bingbing, Yuan, Geng, Wang, Zigeng, Huang, Shaoyi, Peng, Hongwu, Behnam, Payman, Wen, Wujie, Liu, Hang, Ding, Caiwen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ARAS: An Adaptive Low-Cost ReRAM-Based Accelerator for DNNs
by: Sabri, Mohammad, et al.
Published: (2024)
by: Sabri, Mohammad, et al.
Published: (2024)
Online Soft Error Tolerance in ReRAM Crossbars for Deep Learning Accelerators
by: Khezeli, Benyamin, et al.
Published: (2024)
by: Khezeli, Benyamin, et al.
Published: (2024)
All-in-Memory Stochastic Computing using ReRAM
by: de Lima, João Paulo C., et al.
Published: (2025)
by: de Lima, João Paulo C., et al.
Published: (2025)
HURRY: Highly Utilized, Reconfigurable ReRAM-based In-situ Accelerator with Multifunctionality
by: Shin, Hery, et al.
Published: (2024)
by: Shin, Hery, et al.
Published: (2024)
Hamun: An Approximate Computation Method to Prolong the Lifespan of ReRAM-Based Accelerators
by: Sabri, Mohammad, et al.
Published: (2025)
by: Sabri, Mohammad, et al.
Published: (2025)
ReCross: Efficient Embedding Reduction Scheme for In-Memory Computing using ReRAM-Based Crossbar
by: Lai, Yu-Hong, et al.
Published: (2025)
by: Lai, Yu-Hong, et al.
Published: (2025)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
by: Zhang, Chenguang, et al.
Published: (2024)
by: Zhang, Chenguang, et al.
Published: (2024)
Sensitivity-Aware Mixed-Precision Quantization for ReRAM-based Computing-in-Memory
by: Chen, Guan-Cheng, et al.
Published: (2025)
by: Chen, Guan-Cheng, et al.
Published: (2025)
FARe: Fault-Aware GNN Training on ReRAM-based PIM Accelerators
by: Dhingra, Pratyush, et al.
Published: (2024)
by: Dhingra, Pratyush, et al.
Published: (2024)
A Fully Hardware Implemented Accelerator Design in ReRAM Analog Computing without ADCs
by: Dang, Peng, et al.
Published: (2024)
by: Dang, Peng, et al.
Published: (2024)
Stuck-at Faults in ReRAM Neuromorphic Circuit Array and their Correction through Machine Learning
by: Sawal, Vedant, et al.
Published: (2024)
by: Sawal, Vedant, et al.
Published: (2024)
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
by: Shao, Kunming, et al.
Published: (2025)
by: Shao, Kunming, et al.
Published: (2025)
Pointer: An Energy-Efficient ReRAM-based Point Cloud Recognition Accelerator with Inter-layer and Intra-layer Optimizations
by: Zhang, Qijun, et al.
Published: (2024)
by: Zhang, Qijun, et al.
Published: (2024)
4T2R X-ReRAM CiM Array for Variation-tolerant, Low-power, Massively Parallel MAC Operation
by: Kihara, Fuyuki, et al.
Published: (2025)
by: Kihara, Fuyuki, et al.
Published: (2025)
All-in-One Analog AI Hardware: On-Chip Training and Inference with Conductive-Metal-Oxide/HfOx ReRAM Devices
by: Falcone, Donato Francesco, et al.
Published: (2025)
by: Falcone, Donato Francesco, et al.
Published: (2025)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
by: Dong, Pingcheng, et al.
Published: (2026)
by: Dong, Pingcheng, et al.
Published: (2026)
ReTern: Exploiting Natural Redundancy and Sign Transformations for Enhanced Fault Tolerance in Compute-in-Memory based Ternary LLMs
by: Malhotra, Akul, et al.
Published: (2025)
by: Malhotra, Akul, et al.
Published: (2025)
TL-nvSRAM-CIM: Ultra-High-Density Three-Level ReRAM-Assisted Computing-in-nvSRAM with DC-Power Free Restore and Ternary MAC Operations
by: Wang, Dengfeng, et al.
Published: (2023)
by: Wang, Dengfeng, et al.
Published: (2023)
Combining Fault Tolerance Techniques and COTS SoC Accelerators for Payload Processing in Space
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
LLM-VeriPPA: Power, Performance, and Area Optimization aware Verilog Code Generation with Large Language Models
by: Thorat, Kiran, et al.
Published: (2025)
by: Thorat, Kiran, et al.
Published: (2025)
DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference
by: Wen, Jinqi, et al.
Published: (2026)
by: Wen, Jinqi, et al.
Published: (2026)
An ECC-based Fault Tolerance Approach for DNNs
by: Raji, Mohsen, et al.
Published: (2025)
by: Raji, Mohsen, et al.
Published: (2025)
Oobleck: Low-Compromise Design for Fault Tolerant Accelerators
by: Wilks, Guy, et al.
Published: (2025)
by: Wilks, Guy, et al.
Published: (2025)
Weight Transformations in Bit-Sliced Crossbar Arrays for Fault Tolerant Computing-in-Memory: Design Techniques and Evaluation Framework
by: Malhotra, Akul, et al.
Published: (2025)
by: Malhotra, Akul, et al.
Published: (2025)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault Tolerance
by: Xie, Tong, et al.
Published: (2025)
by: Xie, Tong, et al.
Published: (2025)
Testing and Fault Tolerance Techniques for CNT-Based FPGAs
by: Lu, Siyuan, et al.
Published: (2025)
by: Lu, Siyuan, et al.
Published: (2025)
A Reconfigurable Approximate Computing RISC-V Platform for Fault-Tolerant Applications
by: Delavari, Arvin, et al.
Published: (2024)
by: Delavari, Arvin, et al.
Published: (2024)
Reuse Detector: Improving the Management of STT-RAM SLLCs
by: RodrÍguez-RodrÍguez, Roberto, et al.
Published: (2024)
by: RodrÍguez-RodrÍguez, Roberto, et al.
Published: (2024)
Special Session: Sustainable Deployment of Deep Neural Networks on Non-Volatile Compute-in-Memory Accelerators
by: Qin, Yifan, et al.
Published: (2025)
by: Qin, Yifan, et al.
Published: (2025)
DreamRAM: A Fine-Grained Configurable Design Space Modeling Tool for Custom 3D Die-Stacked DRAM
by: Cai, Victor, et al.
Published: (2025)
by: Cai, Victor, et al.
Published: (2025)
Synchronization for Fault-Tolerant Quantum Computers
by: Maurya, Satvik, et al.
Published: (2025)
by: Maurya, Satvik, et al.
Published: (2025)
Voxel-CIM: An Efficient Compute-in-Memory Accelerator for Voxel-based Point Cloud Neural Networks
by: Lin, Xipeng, et al.
Published: (2024)
by: Lin, Xipeng, et al.
Published: (2024)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
HOPE: Holistic STT-RAM Architecture Exploration Framework for Future Cross-Platform Analysis
by: SeyedFaraji, Saeed, et al.
Published: (2024)
by: SeyedFaraji, Saeed, et al.
Published: (2024)
Who Checks the Checker? Enhancing Component-level Architectural SEU Fault Tolerance for End-to-End SoC Protection
by: Rogenmoser, Michael, et al.
Published: (2026)
by: Rogenmoser, Michael, et al.
Published: (2026)
Monad: Towards Cost-effective Specialization for Chiplet-based Spatial Accelerators
by: Hao, Xiaochen, et al.
Published: (2023)
by: Hao, Xiaochen, et al.
Published: (2023)
LinkBo: An Adaptive Single-Wire, Low-Latency, and Fault-Tolerant Communications Interface for Variable-Distance Chip-to-Chip Systems
by: Ye, Bochen, et al.
Published: (2025)
by: Ye, Bochen, et al.
Published: (2025)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
by: Cherezova, Natalia, et al.
Published: (2025)
by: Cherezova, Natalia, et al.
Published: (2025)
Similar Items
-
ARAS: An Adaptive Low-Cost ReRAM-Based Accelerator for DNNs
by: Sabri, Mohammad, et al.
Published: (2024) -
Online Soft Error Tolerance in ReRAM Crossbars for Deep Learning Accelerators
by: Khezeli, Benyamin, et al.
Published: (2024) -
All-in-Memory Stochastic Computing using ReRAM
by: de Lima, João Paulo C., et al.
Published: (2025) -
HURRY: Highly Utilized, Reconfigurable ReRAM-based In-situ Accelerator with Multifunctionality
by: Shin, Hery, et al.
Published: (2024) -
Hamun: An Approximate Computation Method to Prolong the Lifespan of ReRAM-Based Accelerators
by: Sabri, Mohammad, et al.
Published: (2025)