Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yuyang, Cai, Jingjing, Ren, Jiayi, Zhou, Peng, Zhang, Danyang, Du, Yin, Li, Shijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
von: Mukherjee, Dipayan, et al.
Veröffentlicht: (2024)
von: Mukherjee, Dipayan, et al.
Veröffentlicht: (2024)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024)
von: He, Jiaao, et al.
Veröffentlicht: (2024)
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
von: Kharma, Sami, et al.
Veröffentlicht: (2026)
von: Kharma, Sami, et al.
Veröffentlicht: (2026)
Morlet wavelet transform using attenuated sliding Fourier transform and kernel integral for graphic processing unit
von: Yamashita, Yukihiko, et al.
Veröffentlicht: (2021)
von: Yamashita, Yukihiko, et al.
Veröffentlicht: (2021)
Intent-driven scheduling of backup jobs
von: Dutta, Souvik, et al.
Veröffentlicht: (2024)
von: Dutta, Souvik, et al.
Veröffentlicht: (2024)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
Serverless Cold Starts and Where to Find Them
von: Joosen, Artjom, et al.
Veröffentlicht: (2024)
von: Joosen, Artjom, et al.
Veröffentlicht: (2024)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
von: Jallouli, Chaimae, et al.
Veröffentlicht: (2026)
von: Jallouli, Chaimae, et al.
Veröffentlicht: (2026)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
von: Yuan, Renzhong, et al.
Veröffentlicht: (2026)
von: Yuan, Renzhong, et al.
Veröffentlicht: (2026)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
von: Putra, Jody Almaida
Veröffentlicht: (2026)
von: Putra, Jody Almaida
Veröffentlicht: (2026)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
von: Salazar, José Daniel Montoya
Veröffentlicht: (2026)
von: Salazar, José Daniel Montoya
Veröffentlicht: (2026)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
von: Heidari, Sina, et al.
Veröffentlicht: (2026)
von: Heidari, Sina, et al.
Veröffentlicht: (2026)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)
von: Kolluru, Saicharan
Veröffentlicht: (2025)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
von: Dutilleul, Alban, et al.
Veröffentlicht: (2024)
von: Dutilleul, Alban, et al.
Veröffentlicht: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
von: Miletto, Marcelo Cogo, et al.
Veröffentlicht: (2024)
von: Miletto, Marcelo Cogo, et al.
Veröffentlicht: (2024)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
von: Maillou, Vincent, et al.
Veröffentlicht: (2025)
von: Maillou, Vincent, et al.
Veröffentlicht: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference
von: Nguyen, An Xuan
Veröffentlicht: (2026)
von: Nguyen, An Xuan
Veröffentlicht: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
von: Reinhard, Anton, et al.
Veröffentlicht: (2025)
von: Reinhard, Anton, et al.
Veröffentlicht: (2025)
An Online Probabilistic Distributed Tracing System
von: Toslali, M., et al.
Veröffentlicht: (2024)
von: Toslali, M., et al.
Veröffentlicht: (2024)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
von: Mukherjee, Dipayan, et al.
Veröffentlicht: (2024) -
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024) -
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
von: Kharma, Sami, et al.
Veröffentlicht: (2026) -
Morlet wavelet transform using attenuated sliding Fourier transform and kernel integral for graphic processing unit
von: Yamashita, Yukihiko, et al.
Veröffentlicht: (2021) -
Intent-driven scheduling of backup jobs
von: Dutta, Souvik, et al.
Veröffentlicht: (2024)