ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault Tolerance
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Tong, Zhao, Jiawang, Wan, Zishen, Zhang, Zuodong, Wang, Yuan, Wang, Runsheng, Huang, Ru, Li, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference
by: Wen, Jinqi, et al.
Published: (2026)
by: Wen, Jinqi, et al.
Published: (2026)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
Efficient yet Accurate End-to-End SC Accelerator Design
by: Li, Meng, et al.
Published: (2024)
by: Li, Meng, et al.
Published: (2024)
CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-Design
by: Wan, Zishen, et al.
Published: (2025)
by: Wan, Zishen, et al.
Published: (2025)
FLARE: One-Shot PE-Level Fault Localization in Systolic Arrays via Algebraic Test Vectors
by: Venkatasubramanian, Logashree, et al.
Published: (2026)
by: Venkatasubramanian, Logashree, et al.
Published: (2026)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
by: Zhong, Linfeng, et al.
Published: (2025)
by: Zhong, Linfeng, et al.
Published: (2025)
Testing and Fault Tolerance Techniques for CNT-Based FPGAs
by: Lu, Siyuan, et al.
Published: (2025)
by: Lu, Siyuan, et al.
Published: (2025)
An ECC-based Fault Tolerance Approach for DNNs
by: Raji, Mohsen, et al.
Published: (2025)
by: Raji, Mohsen, et al.
Published: (2025)
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
by: Zhu, Jingchen, et al.
Published: (2024)
by: Zhu, Jingchen, et al.
Published: (2024)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
by: Cherezova, Natalia, et al.
Published: (2025)
by: Cherezova, Natalia, et al.
Published: (2025)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Oobleck: Low-Compromise Design for Fault Tolerant Accelerators
by: Wilks, Guy, et al.
Published: (2025)
by: Wilks, Guy, et al.
Published: (2025)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
ADDT -- A Digital Twin Framework for Proactive Safety Validation in Autonomous Driving Systems
by: Yu, Bo, et al.
Published: (2025)
by: Yu, Bo, et al.
Published: (2025)
No Redundancy, No Stall: Lightweight Streaming 3D Gaussian Splatting for Real-time Rendering
by: Wei, Linye, et al.
Published: (2025)
by: Wei, Linye, et al.
Published: (2025)
SCALE-Sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis
by: Raj, Ritik, et al.
Published: (2025)
by: Raj, Ritik, et al.
Published: (2025)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
by: Zhao, Chunyuan, et al.
Published: (2025)
by: Zhao, Chunyuan, et al.
Published: (2025)
A Reconfigurable Approximate Computing RISC-V Platform for Fault-Tolerant Applications
by: Delavari, Arvin, et al.
Published: (2024)
by: Delavari, Arvin, et al.
Published: (2024)
LayoutCopilot: An LLM-powered Multi-agent Collaborative Framework for Interactive Analog Layout Design
by: Liu, Bingyang, et al.
Published: (2024)
by: Liu, Bingyang, et al.
Published: (2024)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
by: Xu, Weikai, et al.
Published: (2026)
by: Xu, Weikai, et al.
Published: (2026)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
by: Wang, Peipei, et al.
Published: (2025)
by: Wang, Peipei, et al.
Published: (2025)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
by: Duan, Cenlin, et al.
Published: (2025)
by: Duan, Cenlin, et al.
Published: (2025)
Combining Fault Tolerance Techniques and COTS SoC Accelerators for Payload Processing in Space
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
ENFOR-SA: End-to-end Cross-layer Transient Fault Injector for Efficient and Accurate DNN Reliability Assessment on Systolic Arrays
by: Tonetto, Rafael Billig, et al.
Published: (2026)
by: Tonetto, Rafael Billig, et al.
Published: (2026)
Synchronization for Fault-Tolerant Quantum Computers
by: Maurya, Satvik, et al.
Published: (2025)
by: Maurya, Satvik, et al.
Published: (2025)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
by: Fan, Zhenkun, et al.
Published: (2026)
by: Fan, Zhenkun, et al.
Published: (2026)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
SEGA-DCIM: Design Space Exploration-Guided Automatic Digital CIM Compiler with Multiple Precision Support
by: Diao, Haikang, et al.
Published: (2025)
by: Diao, Haikang, et al.
Published: (2025)
Weight Transformations in Bit-Sliced Crossbar Arrays for Fault Tolerant Computing-in-Memory: Design Techniques and Evaluation Framework
by: Malhotra, Akul, et al.
Published: (2025)
by: Malhotra, Akul, et al.
Published: (2025)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
by: Lin, Chenqi, et al.
Published: (2025)
by: Lin, Chenqi, et al.
Published: (2025)
CellE: Automated Standard Cell Library Extension via Equality Saturation
by: Ren, Yi, et al.
Published: (2026)
by: Ren, Yi, et al.
Published: (2026)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
by: Zhu, Zhantong, et al.
Published: (2025)
by: Zhu, Zhantong, et al.
Published: (2025)
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
by: Li, Bingbing, et al.
Published: (2024)
by: Li, Bingbing, et al.
Published: (2024)
LSQCA: Resource-Efficient Load/Store Architecture for Limited-Scale Fault-Tolerant Quantum Computing
by: Kobori, Takumi, et al.
Published: (2024)
by: Kobori, Takumi, et al.
Published: (2024)
Similar Items
-
DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference
by: Wen, Jinqi, et al.
Published: (2026) -
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
by: Li, Meng, et al.
Published: (2026) -
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
by: Xie, Tong, et al.
Published: (2026) -
CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems
by: Xie, Tong, et al.
Published: (2026) -
Efficient yet Accurate End-to-End SC Accelerator Design
by: Li, Meng, et al.
Published: (2024)