LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Tyagi, Abhishek, Hukerikar, Saurabh, Saxena, Nirmal, Huang, Yanxiang, Shirvani, Philip, Tung, Chung-Hsuan, Zhu, Yuhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
by: Tung, Chung-Hsuan, et al.
Published: (2026)
by: Tung, Chung-Hsuan, et al.
Published: (2026)
On the Vulnerability of FHE Computation to Silent Data Corruption
by: Mu, Jianan, et al.
Published: (2026)
by: Mu, Jianan, et al.
Published: (2026)
Characterizing Soft-Error Resiliency in Arm's Ethos-U55 Embedded Machine Learning Accelerator
by: Tyagi, Abhishek, et al.
Published: (2024)
by: Tyagi, Abhishek, et al.
Published: (2024)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026)
by: Chung, Euijun, et al.
Published: (2026)
ITHICA: Intra-Thread Instruction Checking Approach for Defect-Induced Silent Data Corruptions
by: Vavelidou, Ioanna, et al.
Published: (2026)
by: Vavelidou, Ioanna, et al.
Published: (2026)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Efficient LLM inference solution on Intel GPU
by: Wu, Hui, et al.
Published: (2023)
by: Wu, Hui, et al.
Published: (2023)
Analysis of LLM Vulnerability to GPU Soft Errors: An Instruction-Level Fault Injection Study
by: Chai, Duo, et al.
Published: (2025)
by: Chai, Duo, et al.
Published: (2025)
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
by: Alama, Youssef A. Ait, et al.
Published: (2024)
by: Alama, Youssef A. Ait, et al.
Published: (2024)
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
by: Gundawar, Ayush, et al.
Published: (2024)
by: Gundawar, Ayush, et al.
Published: (2024)
Not All Faults Are Equal: Transient-Fault Sensitivity Characterization of an Open-Source RISC-V Vector Cluster
by: Cai, Maoyuan, et al.
Published: (2026)
by: Cai, Maoyuan, et al.
Published: (2026)
A Spatio-Temporal Graph Neural Networks Approach for Predicting Silent Data Corruption inducing Circuit-Level Faults
by: Wei, Shaoqi, et al.
Published: (2025)
by: Wei, Shaoqi, et al.
Published: (2025)
Large Language Model (LLM) for Standard Cell Layout Design Optimization
by: Ho, Chia-Tung, et al.
Published: (2024)
by: Ho, Chia-Tung, et al.
Published: (2024)
Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing
by: Mitra, Subhasish, et al.
Published: (2025)
by: Mitra, Subhasish, et al.
Published: (2025)
Assessing the Performance of Stateful Logic in 1-Selector-1-RRAM Crossbar Arrays
by: Tyagi, Arjun, et al.
Published: (2024)
by: Tyagi, Arjun, et al.
Published: (2024)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
by: Zhang, Tao, et al.
Published: (2026)
by: Zhang, Tao, et al.
Published: (2026)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
by: Latif, Imran, et al.
Published: (2024)
by: Latif, Imran, et al.
Published: (2024)
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
by: Bai, Kangbo, et al.
Published: (2026)
by: Bai, Kangbo, et al.
Published: (2026)
Vulnerabilities in Partial TEE-Shielded LLM Inference with Precomputed Noise
by: Saini, Abhishek, et al.
Published: (2026)
by: Saini, Abhishek, et al.
Published: (2026)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
by: Pu, Huanzhi, et al.
Published: (2025)
by: Pu, Huanzhi, et al.
Published: (2025)
GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference
by: Chen, Guoci, et al.
Published: (2026)
by: Chen, Guoci, et al.
Published: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
Thales: Formulating and Estimating Architectural Vulnerability Factors for DNN Accelerators
by: Tyagi, Abhishek, et al.
Published: (2022)
by: Tyagi, Abhishek, et al.
Published: (2022)
OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation
by: Liu, Shang, et al.
Published: (2025)
by: Liu, Shang, et al.
Published: (2025)
VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware Scheduling
by: Lin, Zi-Wei, et al.
Published: (2026)
by: Lin, Zi-Wei, et al.
Published: (2026)
Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
by: Yu, Zhongkai, et al.
Published: (2024)
by: Yu, Zhongkai, et al.
Published: (2024)
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
by: Lin, Zi-Wei, et al.
Published: (2026)
by: Lin, Zi-Wei, et al.
Published: (2026)
Orion: Characterizing and Programming Apple's Neural Engine for LLM Training and Inference
by: Kumaresan, Ramchand
Published: (2026)
by: Kumaresan, Ramchand
Published: (2026)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
by: Mamo, Oteo, et al.
Published: (2026)
by: Mamo, Oteo, et al.
Published: (2026)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Honest to a Fault: Root-Causing Fault Attacks with Pre-Silicon RISC Pipeline Characterization
by: Malik, Arsalan Ali, et al.
Published: (2025)
by: Malik, Arsalan Ali, et al.
Published: (2025)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
by: Anik, Shafayat Mowla, et al.
Published: (2026)
by: Anik, Shafayat Mowla, et al.
Published: (2026)
HLSPilot: LLM-based High-Level Synthesis
by: Xiong, Chenwei, et al.
Published: (2024)
by: Xiong, Chenwei, et al.
Published: (2024)
LIMINAL: Exploring The Frontiers of LLM Decode Performance
by: Davies, Michael, et al.
Published: (2025)
by: Davies, Michael, et al.
Published: (2025)
DRC-Coder: Automated DRC Checker Code Generation Using LLM Autonomous Agent
by: Chang, Chen-Chia, et al.
Published: (2024)
by: Chang, Chen-Chia, et al.
Published: (2024)
RoboGPU: Accelerating GPU Collision Detection for Robotics
by: Liu, Lufei, et al.
Published: (2026)
by: Liu, Lufei, et al.
Published: (2026)
Design of a GPU with Heterogeneous Cores for Graphics
by: Tomás, Aurora, et al.
Published: (2026)
by: Tomás, Aurora, et al.
Published: (2026)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
by: Luo, Weile, et al.
Published: (2024)
by: Luo, Weile, et al.
Published: (2024)
Similar Items
-
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
by: Tung, Chung-Hsuan, et al.
Published: (2026) -
On the Vulnerability of FHE Computation to Silent Data Corruption
by: Mu, Jianan, et al.
Published: (2026) -
Characterizing Soft-Error Resiliency in Arm's Ethos-U55 Embedded Machine Learning Accelerator
by: Tyagi, Abhishek, et al.
Published: (2024) -
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026) -
ITHICA: Intra-Thread Instruction Checking Approach for Defect-Induced Silent Data Corruptions
by: Vavelidou, Ioanna, et al.
Published: (2026)