Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Le, Hieu Xuan, Goh, Benjamin, Tang, Quy Anh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
por: Wei, Zhipeng, et al.
Publicado: (2024)
por: Wei, Zhipeng, et al.
Publicado: (2024)
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
por: Maloyan, Narek, et al.
Publicado: (2025)
por: Maloyan, Narek, et al.
Publicado: (2025)
RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
por: Dang, Quy-Anh, et al.
Publicado: (2025)
por: Dang, Quy-Anh, et al.
Publicado: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
por: Maloyan, Narek, et al.
Publicado: (2025)
por: Maloyan, Narek, et al.
Publicado: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
por: Nguyen, Trong-Hieu, et al.
Publicado: (2024)
por: Nguyen, Trong-Hieu, et al.
Publicado: (2024)
Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR
por: Dang, Quy-Anh, et al.
Publicado: (2026)
por: Dang, Quy-Anh, et al.
Publicado: (2026)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
por: Bellibatlu, Rohith Reddy, et al.
Publicado: (2026)
por: Bellibatlu, Rohith Reddy, et al.
Publicado: (2026)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
por: Mu, Wenhan, et al.
Publicado: (2025)
por: Mu, Wenhan, et al.
Publicado: (2025)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
por: Dang, Quy-Anh, et al.
Publicado: (2025)
por: Dang, Quy-Anh, et al.
Publicado: (2025)
RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models
por: Dang, Quy-Anh, et al.
Publicado: (2026)
por: Dang, Quy-Anh, et al.
Publicado: (2026)
Defense against Prompt Injection Attacks via Mixture of Encodings
por: Zhang, Ruiyi, et al.
Publicado: (2025)
por: Zhang, Ruiyi, et al.
Publicado: (2025)
Mixture-of-Personas Language Models for Population Simulation
por: Bui, Ngoc, et al.
Publicado: (2025)
por: Bui, Ngoc, et al.
Publicado: (2025)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
por: Raina, Vyas, et al.
Publicado: (2024)
por: Raina, Vyas, et al.
Publicado: (2024)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
por: Zhao, Shuai, et al.
Publicado: (2023)
por: Zhao, Shuai, et al.
Publicado: (2023)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
por: Wei, Hui, et al.
Publicado: (2024)
por: Wei, Hui, et al.
Publicado: (2024)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
por: Zhang, Wentao, et al.
Publicado: (2025)
por: Zhang, Wentao, et al.
Publicado: (2025)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
por: Tang, Zhenwei, et al.
Publicado: (2026)
por: Tang, Zhenwei, et al.
Publicado: (2026)
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
por: Wu, Junjie, et al.
Publicado: (2026)
por: Wu, Junjie, et al.
Publicado: (2026)
Prompting a Weighting Mechanism into LLM-as-a-Judge in Two-Step: A Case Study
por: Xie, Wenwen, et al.
Publicado: (2025)
por: Xie, Wenwen, et al.
Publicado: (2025)
Pragmatic Metacognitive Prompting Improves LLM Performance on Sarcasm Detection
por: Lee, Joshua, et al.
Publicado: (2024)
por: Lee, Joshua, et al.
Publicado: (2024)
Mixture-of-Instructions: Aligning Large Language Models via Mixture Prompting
por: Xu, Bowen, et al.
Publicado: (2024)
por: Xu, Bowen, et al.
Publicado: (2024)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
por: Yang, Bo, et al.
Publicado: (2026)
por: Yang, Bo, et al.
Publicado: (2026)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
por: Huang, Hui, et al.
Publicado: (2024)
por: Huang, Hui, et al.
Publicado: (2024)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
por: Marioriyad, Arash, et al.
Publicado: (2025)
por: Marioriyad, Arash, et al.
Publicado: (2025)
Think Twice Before You Judge: Mixture of Dual Reasoning Experts for Multimodal Sarcasm Detection
por: Jana, Soumyadeep, et al.
Publicado: (2025)
por: Jana, Soumyadeep, et al.
Publicado: (2025)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
por: Ding, Ruomeng, et al.
Publicado: (2026)
por: Ding, Ruomeng, et al.
Publicado: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
por: Belmadani, Ikram, et al.
Publicado: (2026)
por: Belmadani, Ikram, et al.
Publicado: (2026)
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
por: Pham, Loc, et al.
Publicado: (2026)
por: Pham, Loc, et al.
Publicado: (2026)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
por: Zhang, Xinran
Publicado: (2026)
por: Zhang, Xinran
Publicado: (2026)
Can LLM be a Personalized Judge?
por: Dong, Yijiang River, et al.
Publicado: (2024)
por: Dong, Yijiang River, et al.
Publicado: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
por: Shi, Lin, et al.
Publicado: (2024)
por: Shi, Lin, et al.
Publicado: (2024)
XMainframe: A Large Language Model for Mainframe Modernization
por: Dau, Anh T. V., et al.
Publicado: (2024)
por: Dau, Anh T. V., et al.
Publicado: (2024)
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
por: Faisal, Fahim, et al.
Publicado: (2024)
por: Faisal, Fahim, et al.
Publicado: (2024)
EnsemJudge: Enhancing Reliability in Chinese LLM-Generated Text Detection through Diverse Model Ensembles
por: Wang, Zhuoshang, et al.
Publicado: (2026)
por: Wang, Zhuoshang, et al.
Publicado: (2026)
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
por: Hoang, Nhat M., et al.
Publicado: (2024)
por: Hoang, Nhat M., et al.
Publicado: (2024)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
por: Son, Guijin, et al.
Publicado: (2024)
por: Son, Guijin, et al.
Publicado: (2024)
Investigating Recent Large Language Models for Vietnamese Machine Reading Comprehension
por: Nguyen, Anh Duc, et al.
Publicado: (2025)
por: Nguyen, Anh Duc, et al.
Publicado: (2025)
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
por: Yang, Jingbo, et al.
Publicado: (2026)
por: Yang, Jingbo, et al.
Publicado: (2026)
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
por: Zhang, Qiyuan, et al.
Publicado: (2025)
por: Zhang, Qiyuan, et al.
Publicado: (2025)
How Interpretable are Reasoning Explanations from Prompting Large Language Models?
por: Yeo, Wei Jie, et al.
Publicado: (2024)
por: Yeo, Wei Jie, et al.
Publicado: (2024)
Ejemplares similares
-
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
por: Wei, Zhipeng, et al.
Publicado: (2024) -
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
por: Maloyan, Narek, et al.
Publicado: (2025) -
RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
por: Dang, Quy-Anh, et al.
Publicado: (2025) -
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
por: Maloyan, Narek, et al.
Publicado: (2025) -
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
por: Nguyen, Trong-Hieu, et al.
Publicado: (2024)