GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Xueyi, Zhou, Zhuoneng, Liu, Zitao, Wu, Yongdong
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914589323558912
author Li, Xueyi
Zhou, Zhuoneng
Liu, Zitao
Wu, Yongdong
author_facet Li, Xueyi
Zhou, Zhuoneng
Liu, Zitao
Wu, Yongdong
contents Large language models (LLMs) are increasingly deployed as educational agents for automatic short answer grading (ASAG) in real-world educational environments, significantly boosting assessment efficiency and scalability. However, when these grading agents operate ``in the wild'', their vulnerability to adversarial manipulation raises critical concerns about agent security and trustworthiness. In this paper, we introduce GradingAttack, a fine-grained adversarial attack framework that systematically evaluates the security vulnerabilities of LLM based educational grading agents. Specifically, we design token-level and prompt-level attack strategies that manipulate agent grading outcomes while maintaining high stealth, exposing fundamental weaknesses in current agent deployments. Experiments on multiple datasets demonstrate that both attack strategies effectively compromise grading agents, with prompt-level attacks achieving higher success rates and token-level attacks exhibiting superior stealth capability. Our findings reveal that current LLM based educational agents lack robust defenses against adversarial attacks, underscoring the urgent need for developing secure and trustworthy agent systems for critical educational applications.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00979
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
Li, Xueyi
Zhou, Zhuoneng
Liu, Zitao
Wu, Yongdong
Cryptography and Security
Artificial Intelligence
Computation and Language
Large language models (LLMs) are increasingly deployed as educational agents for automatic short answer grading (ASAG) in real-world educational environments, significantly boosting assessment efficiency and scalability. However, when these grading agents operate ``in the wild'', their vulnerability to adversarial manipulation raises critical concerns about agent security and trustworthiness. In this paper, we introduce GradingAttack, a fine-grained adversarial attack framework that systematically evaluates the security vulnerabilities of LLM based educational grading agents. Specifically, we design token-level and prompt-level attack strategies that manipulate agent grading outcomes while maintaining high stealth, exposing fundamental weaknesses in current agent deployments. Experiments on multiple datasets demonstrate that both attack strategies effectively compromise grading agents, with prompt-level attacks achieving higher success rates and token-level attacks exhibiting superior stealth capability. Our findings reveal that current LLM based educational agents lack robust defenses against adversarial attacks, underscoring the urgent need for developing secure and trustworthy agent systems for critical educational applications.
title GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.00979