MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Shabgahi, Soheil Zibakhsh, Jandali, Yaman, Koushanfar, Farinaz |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
di: Jandali, Yaman, et al.
Pubblicazione: (2025)
di: Jandali, Yaman, et al.
Pubblicazione: (2025)
Trojan Cleansing with Neural Collapse
di: Gu, Xihe, et al.
Pubblicazione: (2024)
di: Gu, Xihe, et al.
Pubblicazione: (2024)
Props for Machine-Learning Security
di: Juels, Ari, et al.
Pubblicazione: (2024)
di: Juels, Ari, et al.
Pubblicazione: (2024)
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2026)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2026)
LayerCollapse: Adaptive compression of neural networks
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2025)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2025)
Magmaw: Modality-Agnostic Adversarial Attacks on Machine Learning-Based Wireless Communication Systems
di: Chang, Jung-Woo, et al.
Pubblicazione: (2023)
di: Chang, Jung-Woo, et al.
Pubblicazione: (2023)
AttestLLM: Efficient Attestation Framework for Billion-scale On-device LLMs
di: Zhang, Ruisi, et al.
Pubblicazione: (2025)
di: Zhang, Ruisi, et al.
Pubblicazione: (2025)
ZORRO: Zero-Knowledge Robustness and Privacy for Split Learning (Full Version)
di: Sheybani, Nojan, et al.
Pubblicazione: (2025)
di: Sheybani, Nojan, et al.
Pubblicazione: (2025)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
di: Spracklen, Joseph, et al.
Pubblicazione: (2026)
di: Spracklen, Joseph, et al.
Pubblicazione: (2026)
Double-Dip: Thwarting Label-Only Membership Inference Attacks with Transfer Learning and Randomization
di: Rajabi, Arezoo, et al.
Pubblicazione: (2024)
di: Rajabi, Arezoo, et al.
Pubblicazione: (2024)
Watermarking Large Language Models and the Generated Content: Opportunities and Challenges
di: Zhang, Ruisi, et al.
Pubblicazione: (2024)
di: Zhang, Ruisi, et al.
Pubblicazione: (2024)
EmMark: Robust Watermarks for IP Protection of Embedded Quantized Large Language Models
di: Zhang, Ruisi, et al.
Pubblicazione: (2024)
di: Zhang, Ruisi, et al.
Pubblicazione: (2024)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
HOACS: Homomorphic Obfuscation Assisted Concealing of Secrets to Thwart Trojan Attacks in COTS Processor
di: Hossain, Tanvir, et al.
Pubblicazione: (2024)
di: Hossain, Tanvir, et al.
Pubblicazione: (2024)
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
di: Wang, Hongtao, et al.
Pubblicazione: (2026)
di: Wang, Hongtao, et al.
Pubblicazione: (2026)
Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging
di: Latibari, Banafsheh Saber, et al.
Pubblicazione: (2025)
di: Latibari, Banafsheh Saber, et al.
Pubblicazione: (2025)
Trojans in Artificial Intelligence (TrojAI) Final Report
di: Reese, Kristopher W., et al.
Pubblicazione: (2026)
di: Reese, Kristopher W., et al.
Pubblicazione: (2026)
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
di: Faruque, Md Omar, et al.
Pubblicazione: (2024)
di: Faruque, Md Omar, et al.
Pubblicazione: (2024)
ICtoken: An NFT for Hardware IP Protection
di: Balla, Shashank, et al.
Pubblicazione: (2024)
di: Balla, Shashank, et al.
Pubblicazione: (2024)
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
di: Saha, Shoumik, et al.
Pubblicazione: (2026)
di: Saha, Shoumik, et al.
Pubblicazione: (2026)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
di: Zhang, Xu, et al.
Pubblicazione: (2025)
di: Zhang, Xu, et al.
Pubblicazione: (2025)
Large Language Models Merging for Enhancing the Link Stealing Attack on Graph Neural Networks
di: Guan, Faqian, et al.
Pubblicazione: (2024)
di: Guan, Faqian, et al.
Pubblicazione: (2024)
Defending Against Beta Poisoning Attacks in Machine Learning Models
di: Gulciftci, Nilufer, et al.
Pubblicazione: (2025)
di: Gulciftci, Nilufer, et al.
Pubblicazione: (2025)
LiveTune: Dynamic Parameter Tuning for Feedback-Driven Optimization
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
Quantum Properties Trojans (QuPTs) for Attacking Quantum Neural Networks
di: Bhowmik, Sounak, et al.
Pubblicazione: (2025)
di: Bhowmik, Sounak, et al.
Pubblicazione: (2025)
X-Guard: Multilingual Guard Agent for Content Moderation
di: Upadhayay, Bibek, et al.
Pubblicazione: (2025)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
di: Li, Siyuan, et al.
Pubblicazione: (2026)
di: Li, Siyuan, et al.
Pubblicazione: (2026)
CoT-Guard: Small Models for Strong Monitoring
di: Diwan, Nirav, et al.
Pubblicazione: (2026)
di: Diwan, Nirav, et al.
Pubblicazione: (2026)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
di: Wang, Haoran, et al.
Pubblicazione: (2023)
di: Wang, Haoran, et al.
Pubblicazione: (2023)
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
di: Das, Debeshee, et al.
Pubblicazione: (2026)
di: Das, Debeshee, et al.
Pubblicazione: (2026)
Adversarial Machine Learning: Attacks, Defenses, and Open Challenges
di: Jha, Pranav K
Pubblicazione: (2025)
di: Jha, Pranav K
Pubblicazione: (2025)
REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
di: Zhang, Ruisi, et al.
Pubblicazione: (2023)
di: Zhang, Ruisi, et al.
Pubblicazione: (2023)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
di: Li, Qinfeng, et al.
Pubblicazione: (2025)
di: Li, Qinfeng, et al.
Pubblicazione: (2025)
Robustness Analysis of Machine Learning Models for IoT Intrusion Detection Under Data Poisoning Attacks
di: Wulnye, Fortunatus Aabangbio, et al.
Pubblicazione: (2026)
di: Wulnye, Fortunatus Aabangbio, et al.
Pubblicazione: (2026)
Zero-Knowledge Proof Frameworks: A Systematic Survey
di: Sheybani, Nojan, et al.
Pubblicazione: (2025)
di: Sheybani, Nojan, et al.
Pubblicazione: (2025)
LoBAM: LoRA-Based Backdoor Attack on Model Merging
di: Yin, Ming, et al.
Pubblicazione: (2024)
di: Yin, Ming, et al.
Pubblicazione: (2024)
SWIFT: Semantic Watermarking for Image Forgery Thwarting
di: Evennou, Gautier, et al.
Pubblicazione: (2024)
di: Evennou, Gautier, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
di: Jandali, Yaman, et al.
Pubblicazione: (2025) -
Trojan Cleansing with Neural Collapse
di: Gu, Xihe, et al.
Pubblicazione: (2024) -
Props for Machine-Learning Security
di: Juels, Ari, et al.
Pubblicazione: (2024) -
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2026) -
LayerCollapse: Adaptive compression of neural networks
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)