Hu, X., Chen, P., & Ho, T. (2024). Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes.
Chicago-Zitierstil (17. Ausg.)Hu, Xiaomeng, Pin-Yu Chen, und Tsung-Yi Ho. Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes. 2024.
MLA-Zitierstil (9. Ausg.)Hu, Xiaomeng, et al. Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes. 2024.
Achtung: Diese Zitate sind unter Umständen nicht zu 100% korrekt.