Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Xianglin, Hooi, Bryan, Deng, Gelei, Zhang, Tianwei, Dong, Jin Song |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
by: Yang, Xianglin, et al.
Published: (2026)
by: Yang, Xianglin, et al.
Published: (2026)
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
by: Ou, Haoran, et al.
Published: (2026)
by: Ou, Haoran, et al.
Published: (2026)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors
by: Lan, Tianwei, et al.
Published: (2025)
by: Lan, Tianwei, et al.
Published: (2025)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
by: Ou, Haoran, et al.
Published: (2025)
by: Ou, Haoran, et al.
Published: (2025)
LLM-enabled Applications Require System-Level Threat Monitoring
by: Zhang, Yedi, et al.
Published: (2026)
by: Zhang, Yedi, et al.
Published: (2026)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
Verification of Bit-Flip Attacks against Quantized Neural Networks
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
by: Shi, Jiawen, et al.
Published: (2024)
by: Shi, Jiawen, et al.
Published: (2024)
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
by: Zeng, Qirun, et al.
Published: (2025)
by: Zeng, Qirun, et al.
Published: (2025)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
by: Cao, Tri, et al.
Published: (2025)
by: Cao, Tri, et al.
Published: (2025)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
by: Ding, Ruomeng, et al.
Published: (2026)
by: Ding, Ruomeng, et al.
Published: (2026)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
by: Lin, Weiran, et al.
Published: (2024)
by: Lin, Weiran, et al.
Published: (2024)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
SPICED: Syntactical Bug and Trojan Pattern Identification in A/MS Circuits using LLM-Enhanced Detection
by: Chaudhuri, Jayeeta, et al.
Published: (2024)
by: Chaudhuri, Jayeeta, et al.
Published: (2024)
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection
by: Kulkarni, Prashant
Published: (2026)
by: Kulkarni, Prashant
Published: (2026)
Frequency maps reveal the correlation between Adversarial Attacks and Implicit Bias
by: Basile, Lorenzo, et al.
Published: (2023)
by: Basile, Lorenzo, et al.
Published: (2023)
Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
by: Zhou, Xingfu, et al.
Published: (2025)
by: Zhou, Xingfu, et al.
Published: (2025)
Multimodal Large Language Models for Phishing Webpage Detection and Identification
by: Lee, Jehyun, et al.
Published: (2024)
by: Lee, Jehyun, et al.
Published: (2024)
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
EAB-FL: Exacerbating Algorithmic Bias through Model Poisoning Attacks in Federated Learning
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
by: Yang, Junxiao, et al.
Published: (2025)
by: Yang, Junxiao, et al.
Published: (2025)
RoBCtrl: Attacking GNN-Based Social Bot Detectors via Reinforced Manipulation of Bots Control Interaction
by: Yang, Yingguang, et al.
Published: (2025)
by: Yang, Yingguang, et al.
Published: (2025)
Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning
by: Shan, Junjie, et al.
Published: (2024)
by: Shan, Junjie, et al.
Published: (2024)
Evaluating Large Language Models for Security Bug Report Prediction
by: Soltaniani, Farnaz, et al.
Published: (2026)
by: Soltaniani, Farnaz, et al.
Published: (2026)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
by: Zhang, Xinkai, et al.
Published: (2026)
by: Zhang, Xinkai, et al.
Published: (2026)
Backdoor Attack on Vertical Federated Graph Neural Network Learning
by: Yang, Jirui, et al.
Published: (2024)
by: Yang, Jirui, et al.
Published: (2024)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
by: Li, Pengcheng, et al.
Published: (2026)
by: Li, Pengcheng, et al.
Published: (2026)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
by: Yue, Murong, et al.
Published: (2025)
by: Yue, Murong, et al.
Published: (2025)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
by: Jiang, Shuli, et al.
Published: (2024)
by: Jiang, Shuli, et al.
Published: (2024)
Similar Items
-
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025) -
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
by: Yang, Xianglin, et al.
Published: (2026) -
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
by: Ou, Haoran, et al.
Published: (2026) -
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025) -
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)