BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
Fuente:
arXiv
Guardado en:
| Autores principales: | Tong, Terry, Wang, Fei, Zhao, Zhe, Chen, Muhao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2023)
por: Xu, Jiashu, et al.
Publicado: (2023)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
por: Tong, Terry, et al.
Publicado: (2024)
por: Tong, Terry, et al.
Publicado: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
por: Liu, Qin, et al.
Publicado: (2023)
por: Liu, Qin, et al.
Publicado: (2023)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
por: Liu, Qin, et al.
Publicado: (2024)
por: Liu, Qin, et al.
Publicado: (2024)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
por: Ding, Ruomeng, et al.
Publicado: (2026)
por: Ding, Ruomeng, et al.
Publicado: (2026)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
Exploring Backdoor Vulnerabilities of Chat Models
por: Hao, Yunzhuo, et al.
Publicado: (2024)
por: Hao, Yunzhuo, et al.
Publicado: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
por: Owiredu-Ashley, Harry
Publicado: (2026)
por: Owiredu-Ashley, Harry
Publicado: (2026)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
por: Yang, Wenkai, et al.
Publicado: (2024)
por: Yang, Wenkai, et al.
Publicado: (2024)
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
por: Xue, Jiaqi, et al.
Publicado: (2024)
por: Xue, Jiaqi, et al.
Publicado: (2024)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
por: Hu, Man, et al.
Publicado: (2025)
por: Hu, Man, et al.
Publicado: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
por: Jiang, Peihai, et al.
Publicado: (2025)
por: Jiang, Peihai, et al.
Publicado: (2025)
Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model
por: Wu, Tianyi, et al.
Publicado: (2026)
por: Wu, Tianyi, et al.
Publicado: (2026)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
por: Nawal, Aditya, et al.
Publicado: (2026)
por: Nawal, Aditya, et al.
Publicado: (2026)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
por: Du, Wei, et al.
Publicado: (2023)
por: Du, Wei, et al.
Publicado: (2023)
BadActs: A Universal Backdoor Defense in the Activation Space
por: Yi, Biao, et al.
Publicado: (2024)
por: Yi, Biao, et al.
Publicado: (2024)
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
por: Zeng, Shenglai, et al.
Publicado: (2024)
por: Zeng, Shenglai, et al.
Publicado: (2024)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
por: Zhang, Ruyi, et al.
Publicado: (2026)
por: Zhang, Ruyi, et al.
Publicado: (2026)
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
por: Simoni, Marco, et al.
Publicado: (2025)
por: Simoni, Marco, et al.
Publicado: (2025)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
por: Li, Xueyi, et al.
Publicado: (2026)
por: Li, Xueyi, et al.
Publicado: (2026)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
por: Zhao, Shuai, et al.
Publicado: (2025)
por: Zhao, Shuai, et al.
Publicado: (2025)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
por: Cao, Yuanpu, et al.
Publicado: (2023)
por: Cao, Yuanpu, et al.
Publicado: (2023)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
por: Ji, Wence, et al.
Publicado: (2025)
por: Ji, Wence, et al.
Publicado: (2025)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
por: Wei, Jiali, et al.
Publicado: (2026)
por: Wei, Jiali, et al.
Publicado: (2026)
Instructional Fingerprinting of Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2024)
por: Xu, Jiashu, et al.
Publicado: (2024)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
por: Sahoo, Devanshu, et al.
Publicado: (2025)
por: Sahoo, Devanshu, et al.
Publicado: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
por: Schwartz, Daniel, et al.
Publicado: (2025)
por: Schwartz, Daniel, et al.
Publicado: (2025)
Claim-Guided Textual Backdoor Attack for Practical Applications
por: Song, Minkyoo, et al.
Publicado: (2024)
por: Song, Minkyoo, et al.
Publicado: (2024)
Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities
por: Mouzouni, Charafeddine
Publicado: (2026)
por: Mouzouni, Charafeddine
Publicado: (2026)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
por: Cheng, Pengzhou, et al.
Publicado: (2024)
por: Cheng, Pengzhou, et al.
Publicado: (2024)
PentestJudge: Judging Agent Behavior Against Operational Requirements
por: Caldwell, Shane, et al.
Publicado: (2025)
por: Caldwell, Shane, et al.
Publicado: (2025)
Security in LLM-as-a-Judge: A Comprehensive SoK
por: Masoud, Aiman Al, et al.
Publicado: (2026)
por: Masoud, Aiman Al, et al.
Publicado: (2026)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
por: Shi, Jiawen, et al.
Publicado: (2024)
por: Shi, Jiawen, et al.
Publicado: (2024)
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
por: Liu, Hongyi, et al.
Publicado: (2024)
por: Liu, Hongyi, et al.
Publicado: (2024)
BadEdit: Backdooring large language models by model editing
por: Li, Yanzhou, et al.
Publicado: (2024)
por: Li, Yanzhou, et al.
Publicado: (2024)
Ejemplares similares
-
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2023) -
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
por: Tong, Terry, et al.
Publicado: (2024) -
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
por: Liu, Qin, et al.
Publicado: (2023) -
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
por: Liu, Qin, et al.
Publicado: (2024) -
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
por: Ding, Ruomeng, et al.
Publicado: (2026)