LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Songze, Xu, Chuokun, Wang, Jiaying, Gong, Xueluan, Chen, Chen, Zhang, Jirui, Wang, Jun, Lam, Kwok-Yan, Ji, Shouling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
Hidden Data Privacy Breaches in Federated Learning
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025)
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025)
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
von: Ullah, Saad, et al.
Veröffentlicht: (2023)
von: Ullah, Saad, et al.
Veröffentlicht: (2023)
TrojanDam: Detection-Free Backdoor Defense in Federated Learning through Proactive Model Robustification utilizing OOD Data
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
von: Zhou, Kaiyu, et al.
Veröffentlicht: (2026)
von: Zhou, Kaiyu, et al.
Veröffentlicht: (2026)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
URVFL: Undetectable Data Reconstruction Attack on Vertical Federated Learning
von: Yao, Duanyi, et al.
Veröffentlicht: (2024)
von: Yao, Duanyi, et al.
Veröffentlicht: (2024)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
Attributing and Exploiting Safety Vectors through Global Optimization in Large Language Models
von: Chu, Fengheng, et al.
Veröffentlicht: (2026)
von: Chu, Fengheng, et al.
Veröffentlicht: (2026)
ReCIT: Reconstructing Full Private Data from Gradient in Parameter-Efficient Fine-Tuning of Large Language Models
von: Xie, Jin, et al.
Veröffentlicht: (2025)
von: Xie, Jin, et al.
Veröffentlicht: (2025)
Security in LLM-as-a-Judge: A Comprehensive SoK
von: Masoud, Aiman Al, et al.
Veröffentlicht: (2026)
von: Masoud, Aiman Al, et al.
Veröffentlicht: (2026)
Unveiling the Security Risks of Federated Learning in the Wild: From Research to Practice
von: Chen, Jiahao, et al.
Veröffentlicht: (2026)
von: Chen, Jiahao, et al.
Veröffentlicht: (2026)
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMs
von: Xia, Yifan, et al.
Veröffentlicht: (2024)
von: Xia, Yifan, et al.
Veröffentlicht: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
CIBER: A Comprehensive Benchmark for Security Evaluation of Code Interpreter Agents
von: Ba, Lei, et al.
Veröffentlicht: (2026)
von: Ba, Lei, et al.
Veröffentlicht: (2026)
PentestJudge: Judging Agent Behavior Against Operational Requirements
von: Caldwell, Shane, et al.
Veröffentlicht: (2025)
von: Caldwell, Shane, et al.
Veröffentlicht: (2025)
When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack
von: Sun, Zehan, et al.
Veröffentlicht: (2026)
von: Sun, Zehan, et al.
Veröffentlicht: (2026)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
von: Hönig, Robert, et al.
Veröffentlicht: (2024)
von: Hönig, Robert, et al.
Veröffentlicht: (2024)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
AdvSQLi: Generating Adversarial SQL Injections against Real-world WAF-as-a-service
von: Qu, Zhenqing, et al.
Veröffentlicht: (2024)
von: Qu, Zhenqing, et al.
Veröffentlicht: (2024)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Proactive Detection of Physical Inter-rule Vulnerabilities in IoT Services Using a Deep Learning Approach
von: Huang, Bing, et al.
Veröffentlicht: (2024)
von: Huang, Bing, et al.
Veröffentlicht: (2024)
Uncertainty-Aware, Risk-Adaptive Access Control for Agentic Systems using an LLM-Judged TBAC Model
von: Fleming, Charles, et al.
Veröffentlicht: (2025)
von: Fleming, Charles, et al.
Veröffentlicht: (2025)
PSA: Private Set Alignment for Secure and Collaborative Analytics on Large-Scale Data
von: Wang, Jiabo, et al.
Veröffentlicht: (2024)
von: Wang, Jiabo, et al.
Veröffentlicht: (2024)
Manifoldchain: Maximizing Blockchain Throughput via Bandwidth-Clustered Sharding
von: Che, Chunjiang, et al.
Veröffentlicht: (2024)
von: Che, Chunjiang, et al.
Veröffentlicht: (2024)
Marlin: Knowledge-Driven Analysis of Provenance Graphs for Efficient and Robust Detection of Cyber Attacks
von: Li, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Li, Zhenyuan, et al.
Veröffentlicht: (2024)
A Failure-Free and Efficient Discrete Laplace Distribution for Differential Privacy in MPC
von: Tjuawinata, Ivan, et al.
Veröffentlicht: (2025)
von: Tjuawinata, Ivan, et al.
Veröffentlicht: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
Megatron: Evasive Clean-Label Backdoor Attacks against Vision Transformer
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
An Effective and Resilient Backdoor Attack Framework against Deep Neural Networks and Vision Transformers
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
Watch the Watcher! Backdoor Attacks on Security-Enhancing Diffusion Models
von: Li, Changjiang, et al.
Veröffentlicht: (2024)
von: Li, Changjiang, et al.
Veröffentlicht: (2024)
BackdoorIndicator: Leveraging OOD Data for Proactive Backdoor Detection in Federated Learning
von: Li, Songze, et al.
Veröffentlicht: (2024)
von: Li, Songze, et al.
Veröffentlicht: (2024)
Arbitrary-Threshold Fully Homomorphic Encryption with Lower Complexity
von: Chang, Yijia, et al.
Veröffentlicht: (2025)
von: Chang, Yijia, et al.
Veröffentlicht: (2025)
Guaranteeing Data Privacy in Federated Unlearning with Dynamic User Participation
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
von: Chen, Chen, et al.
Veröffentlicht: (2026) -
Hidden Data Privacy Breaches in Federated Learning
von: Gong, Xueluan, et al.
Veröffentlicht: (2024) -
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025) -
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
von: Gong, Xueluan, et al.
Veröffentlicht: (2024) -
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
von: Ullah, Saad, et al.
Veröffentlicht: (2023)