Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-Backdoors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Graf, Victoria, Liu, Qin, Chen, Muhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
Test-time Backdoor Mitigation for Black-Box Large Language Models with Defensive Demonstrations
von: Mo, Wenjie, et al.
Veröffentlicht: (2023)
von: Mo, Wenjie, et al.
Veröffentlicht: (2023)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
PRA-PoE: Robust Multimodal Alzheimer's Diagnosis with Arbitrary Missing Modalities
von: Yang, Guangqian, et al.
Veröffentlicht: (2026)
von: Yang, Guangqian, et al.
Veröffentlicht: (2026)
Defense Against Syntactic Textual Backdoor Attacks with Token Substitution
von: Li, Xinglin, et al.
Veröffentlicht: (2024)
von: Li, Xinglin, et al.
Veröffentlicht: (2024)
Two Heads are Better than One: Robust Learning Meets Multi-branch Models
von: Zhang, Zongyuan, et al.
Veröffentlicht: (2022)
von: Zhang, Zongyuan, et al.
Veröffentlicht: (2022)
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
von: Li, Jiazheng, et al.
Veröffentlicht: (2025)
von: Li, Jiazheng, et al.
Veröffentlicht: (2025)
Two is Better than One: Efficient Ensemble Defense for Robust and Compact Models
von: Jung, Yoojin, et al.
Veröffentlicht: (2025)
von: Jung, Yoojin, et al.
Veröffentlicht: (2025)
Two Heads Are Better Than One: Integrating Knowledge from Knowledge Graphs and Large Language Models for Entity Alignment
von: Yang, Linyao, et al.
Veröffentlicht: (2024)
von: Yang, Linyao, et al.
Veröffentlicht: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models
von: Chew, Oscar, et al.
Veröffentlicht: (2025)
von: Chew, Oscar, et al.
Veröffentlicht: (2025)
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
BadActs: A Universal Backdoor Defense in the Activation Space
von: Yi, Biao, et al.
Veröffentlicht: (2024)
von: Yi, Biao, et al.
Veröffentlicht: (2024)
The Knowledge Microscope: Features as Better Analytical Lenses than Neurons
von: Chen, Yuheng, et al.
Veröffentlicht: (2025)
von: Chen, Yuheng, et al.
Veröffentlicht: (2025)
Pruning Strategies for Backdoor Defense in LLMs
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
von: Shen, Xinjie, et al.
Veröffentlicht: (2026)
von: Shen, Xinjie, et al.
Veröffentlicht: (2026)
Two Heads Better than One: Dual Degradation Representation for Blind Super-Resolution
von: Yuan, Hsuan, et al.
Veröffentlicht: (2025)
von: Yuan, Hsuan, et al.
Veröffentlicht: (2025)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
von: Patel, Het, et al.
Veröffentlicht: (2025)
von: Patel, Het, et al.
Veröffentlicht: (2025)
Monotonic Paraphrasing Improves Generalization of Language Model Prompting
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
von: Su, Haoyang, et al.
Veröffentlicht: (2024)
von: Su, Haoyang, et al.
Veröffentlicht: (2024)
Two Heads are Actually Better than One: Towards Better Adversarial Robustness via Transduction and Rejection
von: Palumbo, Nils, et al.
Veröffentlicht: (2023)
von: Palumbo, Nils, et al.
Veröffentlicht: (2023)
PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Why LoRA Fails to Forget: Regularized Low-Rank Adaptation Against Backdoors in Language Models
von: Luong, Hoang-Chau, et al.
Veröffentlicht: (2026)
von: Luong, Hoang-Chau, et al.
Veröffentlicht: (2026)
PoE-World: Compositional World Modeling with Products of Programmatic Experts
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2025)
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2025)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Data Defenses Against Large Language Models
von: Agnew, William, et al.
Veröffentlicht: (2024)
von: Agnew, William, et al.
Veröffentlicht: (2024)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
DebugLM: Learning Traceable Training Data Provenance for LLMs
von: Mo, Wenjie Jacky, et al.
Veröffentlicht: (2026)
von: Mo, Wenjie Jacky, et al.
Veröffentlicht: (2026)
Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
von: Dong, Yichen, et al.
Veröffentlicht: (2025)
von: Dong, Yichen, et al.
Veröffentlicht: (2025)
Defensive Dual Masking for Robust Adversarial Defense
von: Yang, Wangli, et al.
Veröffentlicht: (2024)
von: Yang, Wangli, et al.
Veröffentlicht: (2024)
Data-centric NLP Backdoor Defense from the Lens of Memorization
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023) -
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024) -
Test-time Backdoor Mitigation for Black-Box Large Language Models with Defensive Demonstrations
von: Mo, Wenjie, et al.
Veröffentlicht: (2023) -
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024) -
Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)