AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Tung-Ling, Wu, Yuhao, Liu, Hongliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021)
by: Roth, Tom, et al.
Published: (2021)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
by: Maloyan, Narek, et al.
Published: (2025)
by: Maloyan, Narek, et al.
Published: (2025)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
by: Eiras, Francisco, et al.
Published: (2025)
by: Eiras, Francisco, et al.
Published: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
by: Zheng, Rui, et al.
Published: (2024)
by: Zheng, Rui, et al.
Published: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
by: Zhang, Anqi, et al.
Published: (2024)
by: Zhang, Anqi, et al.
Published: (2024)
TFL: Targeted Bit-Flip Attack on Large Language Model
by: Guo, Jingkai, et al.
Published: (2026)
by: Guo, Jingkai, et al.
Published: (2026)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
by: Li, Qizhang, et al.
Published: (2024)
by: Li, Qizhang, et al.
Published: (2024)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
by: Guo, Jingkai, et al.
Published: (2025)
by: Guo, Jingkai, et al.
Published: (2025)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
by: Paulus, Anselm, et al.
Published: (2024)
by: Paulus, Anselm, et al.
Published: (2024)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
by: Bailey, Luke, et al.
Published: (2023)
by: Bailey, Luke, et al.
Published: (2023)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
by: Lan, Wenhao, et al.
Published: (2026)
by: Lan, Wenhao, et al.
Published: (2026)
Hijacking Large Language Models via Adversarial In-Context Learning
by: Zhou, Xiangyu, et al.
Published: (2023)
by: Zhou, Xiangyu, et al.
Published: (2023)
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
by: Zhang, Collin, et al.
Published: (2024)
by: Zhang, Collin, et al.
Published: (2024)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
by: Tang, Haochun, et al.
Published: (2026)
by: Tang, Haochun, et al.
Published: (2026)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models
by: Zhang, Tianchen, et al.
Published: (2024)
by: Zhang, Tianchen, et al.
Published: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
by: Zhu, Sicheng, et al.
Published: (2024)
by: Zhu, Sicheng, et al.
Published: (2024)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
by: Liu, Ruixuan, et al.
Published: (2026)
by: Liu, Ruixuan, et al.
Published: (2026)
How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
by: Mostafa, Ahmed, et al.
Published: (2025)
by: Mostafa, Ahmed, et al.
Published: (2025)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
by: Liu, Hongfu, et al.
Published: (2024)
by: Liu, Hongfu, et al.
Published: (2024)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
by: Biswas, Sajib, et al.
Published: (2025)
by: Biswas, Sajib, et al.
Published: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
LLM Unlearning Should Be Form-Independent
by: Ye, Xiaotian, et al.
Published: (2025)
by: Ye, Xiaotian, et al.
Published: (2025)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
by: Zhang, Fangyuan, et al.
Published: (2024)
by: Zhang, Fangyuan, et al.
Published: (2024)
Graded Suspiciousness of Adversarial Texts to Human
by: Tonni, Shakila Mahjabin, et al.
Published: (2024)
by: Tonni, Shakila Mahjabin, et al.
Published: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Proving membership in LLM pretraining data via data watermarks
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
by: Liu, Xiaoze, et al.
Published: (2025)
by: Liu, Xiaoze, et al.
Published: (2025)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
by: Ding, Ruomeng, et al.
Published: (2026)
by: Ding, Ruomeng, et al.
Published: (2026)
Differentially Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
Saliency Attention and Semantic Similarity-Driven Adversarial Perturbation
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
IDT: Dual-Task Adversarial Attacks for Privacy Protection
by: Faustini, Pedro, et al.
Published: (2024)
by: Faustini, Pedro, et al.
Published: (2024)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
by: More, Yash, et al.
Published: (2024)
by: More, Yash, et al.
Published: (2024)
Similar Items
-
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025) -
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025) -
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021) -
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
by: Maloyan, Narek, et al.
Published: (2025) -
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
by: Eiras, Francisco, et al.
Published: (2025)