RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
Fuente:
arXiv
Saved in:
| Main Authors: | Lv, Bo, Xu, Zhiheng, Xiu, KeDong, Ding, Ruyi, Zheng, Tianhang, Wang, Zhibo, Ren, Kui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
by: Ma, Haiyue, et al.
Published: (2025)
by: Ma, Haiyue, et al.
Published: (2025)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
by: Huang, Wei-Hsing, et al.
Published: (2025)
by: Huang, Wei-Hsing, et al.
Published: (2025)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
by: Ma, Songchen, et al.
Published: (2026)
by: Ma, Songchen, et al.
Published: (2026)
Accelerating Detailed Routing Convergence through Offline Reinforcement Learning
by: Khan, Afsara, et al.
Published: (2025)
by: Khan, Afsara, et al.
Published: (2025)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
by: Luo, Shuqing, et al.
Published: (2026)
by: Luo, Shuqing, et al.
Published: (2026)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
by: Yu, Yanpeng, et al.
Published: (2025)
by: Yu, Yanpeng, et al.
Published: (2025)
NISTT: A Non-Intrusive SystemC-TLM 2.0 Tracing Tool
by: Bosbach, Nils, et al.
Published: (2022)
by: Bosbach, Nils, et al.
Published: (2022)
Learning Cache Coherence Traffic for NoC Routing Design
by: Xiong, Guochu, et al.
Published: (2025)
by: Xiong, Guochu, et al.
Published: (2025)
GR-Evolve: Design-Adaptive Global Routing via LLM-Driven Algorithm Evolution
by: Jafri, Taizun, et al.
Published: (2026)
by: Jafri, Taizun, et al.
Published: (2026)
System-Technology Co-Optimization of Bitline Routing and Bonding Pathways in Monolithic 3D DRAM Architectures
by: Lee, Kiseok, et al.
Published: (2026)
by: Lee, Kiseok, et al.
Published: (2026)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
by: Hwang, Ranggi, et al.
Published: (2023)
by: Hwang, Ranggi, et al.
Published: (2023)
UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
by: Zhou, Zhuoshan, et al.
Published: (2026)
by: Zhou, Zhuoshan, et al.
Published: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
by: Zhang, Qijun, et al.
Published: (2026)
by: Zhang, Qijun, et al.
Published: (2026)
Length-Matching Routing for Programmable Photonic Circuits Using Best-First Strategy
by: Wang, Xiaoke, et al.
Published: (2025)
by: Wang, Xiaoke, et al.
Published: (2025)
A Limits Study of Memory-side Tiering Telemetry
by: Petrucci, Vinicius, et al.
Published: (2025)
by: Petrucci, Vinicius, et al.
Published: (2025)
Secure Multi-Path Routing with All-or-Nothing Transform for Network-on-Chip Architectures
by: Weerasena, Hansika, et al.
Published: (2026)
by: Weerasena, Hansika, et al.
Published: (2026)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Automated SVA Generation with LLMs
by: Fu, Lik Tung, et al.
Published: (2026)
by: Fu, Lik Tung, et al.
Published: (2026)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
by: Kim, Jungwoo, et al.
Published: (2026)
by: Kim, Jungwoo, et al.
Published: (2026)
PHAROS: Pipelined Heterogeneous Accelerators for Real-time Safety-critical Systems With Deadline Compliance
by: Ji, Shixin, et al.
Published: (2026)
by: Ji, Shixin, et al.
Published: (2026)
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures
by: Shende, Omkar B, et al.
Published: (2026)
by: Shende, Omkar B, et al.
Published: (2026)
UVLLM: An Automated Universal RTL Verification Framework using LLMs
by: Hu, Yuchen, et al.
Published: (2024)
by: Hu, Yuchen, et al.
Published: (2024)
Device-Level Optimization Techniques for Solid-State Drives: A Survey
by: Ren, Tianyu, et al.
Published: (2025)
by: Ren, Tianyu, et al.
Published: (2025)
Robust Qubit Mapping Algorithm via Double-Source Optimal Routing on Large Quantum Circuits
by: Cheng, Chin-Yi, et al.
Published: (2022)
by: Cheng, Chin-Yi, et al.
Published: (2022)
FlashMoE: Fast Distributed MoE in a Single Kernel
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
by: Jiang, Aojie, et al.
Published: (2026)
by: Jiang, Aojie, et al.
Published: (2026)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
by: Gao, Hanyuan, et al.
Published: (2026)
by: Gao, Hanyuan, et al.
Published: (2026)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
by: Kyung, Kwanhee, et al.
Published: (2025)
by: Kyung, Kwanhee, et al.
Published: (2025)
Deep Learning-Based Anomaly Detection in Spacecraft Telemetry on Edge Devices
by: Goetze, Christopher, et al.
Published: (2026)
by: Goetze, Christopher, et al.
Published: (2026)
ADDT -- A Digital Twin Framework for Proactive Safety Validation in Autonomous Driving Systems
by: Yu, Bo, et al.
Published: (2025)
by: Yu, Bo, et al.
Published: (2025)
PreRoutGNN for Timing Prediction with Order Preserving Partition: Global Circuit Pre-training, Local Delay Learning and Attentional Cell Modeling
by: Zhong, Ruizhe, et al.
Published: (2024)
by: Zhong, Ruizhe, et al.
Published: (2024)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
by: Pan, Yue, et al.
Published: (2025)
by: Pan, Yue, et al.
Published: (2025)
HTM-EAR: Importance-Preserving Tiered Memory with Hybrid Routing under Saturation
by: Singh, Shubham Kumar
Published: (2026)
by: Singh, Shubham Kumar
Published: (2026)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
by: Ye, Hanchen, et al.
Published: (2025)
by: Ye, Hanchen, et al.
Published: (2025)
Faster Inference of LLMs using FP8 on the Intel Gaudi
by: Lee, Joonhyung, et al.
Published: (2025)
by: Lee, Joonhyung, et al.
Published: (2025)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
by: Xu, Weikai, et al.
Published: (2026)
by: Xu, Weikai, et al.
Published: (2026)
Similar Items
-
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
by: Ma, Haiyue, et al.
Published: (2025) -
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025) -
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
by: Huang, Wei-Hsing, et al.
Published: (2025) -
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
by: Ma, Songchen, et al.
Published: (2026) -
Accelerating Detailed Routing Convergence through Offline Reinforcement Learning
by: Khan, Afsara, et al.
Published: (2025)