Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raina, Vyas, Liusie, Adian, Gales, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
Finetuning LLMs for Comparative Assessment Tasks
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
von: Molenda, Piotr, et al.
Veröffentlicht: (2024)
von: Molenda, Piotr, et al.
Veröffentlicht: (2024)
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
Investigating the Emergent Audio Classification Ability of ASR Foundation Models
von: Ma, Rao, et al.
Veröffentlicht: (2023)
von: Ma, Rao, et al.
Veröffentlicht: (2023)
Zero-shot Audio Topic Reranking using Large Language Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2023)
von: Qian, Mengjie, et al.
Veröffentlicht: (2023)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
von: Gupta, Akash, et al.
Veröffentlicht: (2024)
von: Gupta, Akash, et al.
Veröffentlicht: (2024)
Is it Possible to Modify Text to a Target Readability Level? An Initial Investigation Using Zero-Shot Large Language Models
von: Farajidizaji, Asma, et al.
Veröffentlicht: (2023)
von: Farajidizaji, Asma, et al.
Veröffentlicht: (2023)
Question-Based Retrieval using Atomic Units for Enterprise RAG
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
von: He, Eric, et al.
Veröffentlicht: (2025)
von: He, Eric, et al.
Veröffentlicht: (2025)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
Extreme Miscalibration and the Illusion of Adversarial Robustness
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
An Information-Theoretic Approach to Analyze NLP Classification Tasks
von: Wang, Luran, et al.
Veröffentlicht: (2024)
von: Wang, Luran, et al.
Veröffentlicht: (2024)
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
An Investigation of Prompt Variations for Zero-shot LLM-based Rankers
von: Sun, Shuoqi, et al.
Veröffentlicht: (2024)
von: Sun, Shuoqi, et al.
Veröffentlicht: (2024)
Who can we trust? LLM-as-a-jury for Comparative Assessment
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
von: Lee, Dongryeol, et al.
Veröffentlicht: (2024)
von: Lee, Dongryeol, et al.
Veröffentlicht: (2024)
Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
von: Le, Hieu Xuan, et al.
Veröffentlicht: (2026)
von: Le, Hieu Xuan, et al.
Veröffentlicht: (2026)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
von: Mu, Wenhan, et al.
Veröffentlicht: (2025)
von: Mu, Wenhan, et al.
Veröffentlicht: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
SkillAggregation: Reference-free LLM-Dependent Aggregation
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
Investigating Non-Transitivity in LLM-as-a-Judge
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
von: Wei, Zhipeng, et al.
Veröffentlicht: (2024)
von: Wei, Zhipeng, et al.
Veröffentlicht: (2024)
Improving Zero-shot LLM Re-Ranker with Risk Minimization
von: Yuan, Xiaowei, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaowei, et al.
Veröffentlicht: (2024)
Zero-shot Persuasive Chatbots with LLM-Generated Strategies and Information Retrieval
von: Furumai, Kazuaki, et al.
Veröffentlicht: (2024)
von: Furumai, Kazuaki, et al.
Veröffentlicht: (2024)
STACK: Adversarial Attacks on LLM Safeguard Pipelines
von: McKenzie, Ian R., et al.
Veröffentlicht: (2025)
von: McKenzie, Ian R., et al.
Veröffentlicht: (2025)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
von: Kaneko, Masahiro
Veröffentlicht: (2026)
von: Kaneko, Masahiro
Veröffentlicht: (2026)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
von: Liusie, Adian, et al.
Veröffentlicht: (2024) -
Finetuning LLMs for Comparative Assessment Tasks
von: Raina, Vatsal, et al.
Veröffentlicht: (2024) -
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2023) -
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024) -
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
von: Molenda, Piotr, et al.
Veröffentlicht: (2024)