Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts
Fuente:
arXiv
Salvato in:
| Autore principale: | Siddiky, Md Nurul Absar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
di: Kim, Jaehan, et al.
Pubblicazione: (2025)
di: Kim, Jaehan, et al.
Pubblicazione: (2025)
SecMoE: Communication-Efficient Secure MoE Inference via Select-Then-Compute
di: Shen, Bowen, et al.
Pubblicazione: (2026)
di: Shen, Bowen, et al.
Pubblicazione: (2026)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
di: Lai, Zhenglin, et al.
Pubblicazione: (2025)
di: Lai, Zhenglin, et al.
Pubblicazione: (2025)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
di: Ding, Xuwei, et al.
Pubblicazione: (2026)
di: Ding, Xuwei, et al.
Pubblicazione: (2026)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
di: Jones, Jaylen, et al.
Pubblicazione: (2026)
di: Jones, Jaylen, et al.
Pubblicazione: (2026)
Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis
di: Topol, Zvi
Pubblicazione: (2026)
di: Topol, Zvi
Pubblicazione: (2026)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
di: Wang, Qingyue, et al.
Pubblicazione: (2025)
di: Wang, Qingyue, et al.
Pubblicazione: (2025)
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
di: Langiu, Alessio
Pubblicazione: (2026)
di: Langiu, Alessio
Pubblicazione: (2026)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
Bypassing Prompt Injection Detectors through Evasive Injections
di: Rahman, Md Jahedur, et al.
Pubblicazione: (2026)
di: Rahman, Md Jahedur, et al.
Pubblicazione: (2026)
A Jailbroken GenAI Model Can Cause Substantial Harm: GenAI-powered Applications are Vulnerable to PromptWares
di: Cohen, Stav, et al.
Pubblicazione: (2024)
di: Cohen, Stav, et al.
Pubblicazione: (2024)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
di: Zhang, Junbo, et al.
Pubblicazione: (2026)
di: Zhang, Junbo, et al.
Pubblicazione: (2026)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
di: Lu, Guoxin, et al.
Pubblicazione: (2026)
di: Lu, Guoxin, et al.
Pubblicazione: (2026)
FragBench: Cross-Session Attacks Hidden in Benign-Looking Fragments
di: Mehta, Astha, et al.
Pubblicazione: (2026)
di: Mehta, Astha, et al.
Pubblicazione: (2026)
Releasing Malevolence from Benevolence: The Menace of Benign Data on Machine Unlearning
di: Ma, Binhao, et al.
Pubblicazione: (2024)
di: Ma, Binhao, et al.
Pubblicazione: (2024)
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
di: Hasan, Md. Mehedi, et al.
Pubblicazione: (2025)
di: Hasan, Md. Mehedi, et al.
Pubblicazione: (2025)
Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents
di: Dang, Hung
Pubblicazione: (2026)
di: Dang, Hung
Pubblicazione: (2026)
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
di: Li, Guanlin, et al.
Pubblicazione: (2024)
di: Li, Guanlin, et al.
Pubblicazione: (2024)
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
di: Rahman, Md Hafizur, et al.
Pubblicazione: (2026)
di: Rahman, Md Hafizur, et al.
Pubblicazione: (2026)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
di: Kabir, Md Rysul, et al.
Pubblicazione: (2026)
di: Kabir, Md Rysul, et al.
Pubblicazione: (2026)
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
di: Hossain, Saad, et al.
Pubblicazione: (2026)
di: Hossain, Saad, et al.
Pubblicazione: (2026)
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
di: Xiong, Chen, et al.
Pubblicazione: (2026)
di: Xiong, Chen, et al.
Pubblicazione: (2026)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
di: Liu, Guozhi, et al.
Pubblicazione: (2025)
di: Liu, Guozhi, et al.
Pubblicazione: (2025)
Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles
di: Wang, Zhilong, et al.
Pubblicazione: (2024)
di: Wang, Zhilong, et al.
Pubblicazione: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
di: Yang, Jirui, et al.
Pubblicazione: (2025)
di: Yang, Jirui, et al.
Pubblicazione: (2025)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
di: He, Luxi, et al.
Pubblicazione: (2024)
di: He, Luxi, et al.
Pubblicazione: (2024)
Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization
di: Maljkovic, Igor, et al.
Pubblicazione: (2026)
di: Maljkovic, Igor, et al.
Pubblicazione: (2026)
Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection
di: Wang, Zhilong, et al.
Pubblicazione: (2024)
di: Wang, Zhilong, et al.
Pubblicazione: (2024)
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
di: Liang, Jiacheng, et al.
Pubblicazione: (2026)
di: Liang, Jiacheng, et al.
Pubblicazione: (2026)
Fine-Tuning Small Language Models for Solution-Oriented Windows Event Log Analysis
di: Akhtar, Siraaj, et al.
Pubblicazione: (2026)
di: Akhtar, Siraaj, et al.
Pubblicazione: (2026)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
di: Chen, Shuhao, et al.
Pubblicazione: (2026)
di: Chen, Shuhao, et al.
Pubblicazione: (2026)
Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
di: Yang, Fan
Pubblicazione: (2025)
di: Yang, Fan
Pubblicazione: (2025)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
di: Bhatt, Manish, et al.
Pubblicazione: (2026)
di: Bhatt, Manish, et al.
Pubblicazione: (2026)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
di: Chu, Junjie, et al.
Pubblicazione: (2026)
di: Chu, Junjie, et al.
Pubblicazione: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
Privacy-Preserving LLMs Routing
di: Wu, Xidong, et al.
Pubblicazione: (2026)
di: Wu, Xidong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
di: Kim, Jaehan, et al.
Pubblicazione: (2025) -
SecMoE: Communication-Efficient Secure MoE Inference via Select-Then-Compute
di: Shen, Bowen, et al.
Pubblicazione: (2026) -
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
di: Lai, Zhenglin, et al.
Pubblicazione: (2025) -
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
di: Ding, Xuwei, et al.
Pubblicazione: (2026) -
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
di: Jones, Jaylen, et al.
Pubblicazione: (2026)