Global Challenge for Safe and Secure LLMs Track 1

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jia, Xiaojun, Huang, Yihao, Liu, Yang, Tan, Peng Yan, Yau, Weng Kuan, Mak, Mun-Thye, Sim, Xin Ming, Ng, Wee Siong, Ng, See Kiong, Liu, Hanqing, Zhou, Lifeng, Yan, Huanqian, Sun, Xiaobing, Liu, Wei, Wang, Long, Qian, Yiming, Liu, Yong, Yang, Junxiao, Zhang, Zhexin, Lei, Leqi, Chen, Renmiao, Lu, Yida, Cui, Shiyao, Wang, Zizhou, Li, Shaohua, Wang, Yan, Goh, Rick Siow Mong, Zhen, Liangli, Zhang, Yingjie, Zhao, Zhe
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909399637819392
author Jia, Xiaojun
Huang, Yihao
Liu, Yang
Tan, Peng Yan
Yau, Weng Kuan
Mak, Mun-Thye
Sim, Xin Ming
Ng, Wee Siong
Ng, See Kiong
Liu, Hanqing
Zhou, Lifeng
Yan, Huanqian
Sun, Xiaobing
Liu, Wei
Wang, Long
Qian, Yiming
Liu, Yong
Yang, Junxiao
Zhang, Zhexin
Lei, Leqi
Chen, Renmiao
Lu, Yida
Cui, Shiyao
Wang, Zizhou
Li, Shaohua
Wang, Yan
Goh, Rick Siow Mong
Zhen, Liangli
Zhang, Yingjie
Zhao, Zhe
author_facet Jia, Xiaojun
Huang, Yihao
Liu, Yang
Tan, Peng Yan
Yau, Weng Kuan
Mak, Mun-Thye
Sim, Xin Ming
Ng, Wee Siong
Ng, See Kiong
Liu, Hanqing
Zhou, Lifeng
Yan, Huanqian
Sun, Xiaobing
Liu, Wei
Wang, Long
Qian, Yiming
Liu, Yong
Yang, Junxiao
Zhang, Zhexin
Lei, Leqi
Chen, Renmiao
Lu, Yida
Cui, Shiyao
Wang, Zizhou
Li, Shaohua
Wang, Yan
Goh, Rick Siow Mong
Zhen, Liangli
Zhang, Yingjie
Zhao, Zhe
contents This paper introduces the Global Challenge for Safe and Secure Large Language Models (LLMs), a pioneering initiative organized by AI Singapore (AISG) and the CyberSG R&D Programme Office (CRPO) to foster the development of advanced defense mechanisms against automated jailbreaking attacks. With the increasing integration of LLMs in critical sectors such as healthcare, finance, and public administration, ensuring these models are resilient to adversarial attacks is vital for preventing misuse and upholding ethical standards. This competition focused on two distinct tracks designed to evaluate and enhance the robustness of LLM security frameworks. Track 1 tasked participants with developing automated methods to probe LLM vulnerabilities by eliciting undesirable responses, effectively testing the limits of existing safety protocols within LLMs. Participants were challenged to devise techniques that could bypass content safeguards across a diverse array of scenarios, from offensive language to misinformation and illegal activities. Through this process, Track 1 aimed to deepen the understanding of LLM vulnerabilities and provide insights for creating more resilient models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14502
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Global Challenge for Safe and Secure LLMs Track 1
Jia, Xiaojun
Huang, Yihao
Liu, Yang
Tan, Peng Yan
Yau, Weng Kuan
Mak, Mun-Thye
Sim, Xin Ming
Ng, Wee Siong
Ng, See Kiong
Liu, Hanqing
Zhou, Lifeng
Yan, Huanqian
Sun, Xiaobing
Liu, Wei
Wang, Long
Qian, Yiming
Liu, Yong
Yang, Junxiao
Zhang, Zhexin
Lei, Leqi
Chen, Renmiao
Lu, Yida
Cui, Shiyao
Wang, Zizhou
Li, Shaohua
Wang, Yan
Goh, Rick Siow Mong
Zhen, Liangli
Zhang, Yingjie
Zhao, Zhe
Cryptography and Security
Artificial Intelligence
Computers and Society
This paper introduces the Global Challenge for Safe and Secure Large Language Models (LLMs), a pioneering initiative organized by AI Singapore (AISG) and the CyberSG R&D Programme Office (CRPO) to foster the development of advanced defense mechanisms against automated jailbreaking attacks. With the increasing integration of LLMs in critical sectors such as healthcare, finance, and public administration, ensuring these models are resilient to adversarial attacks is vital for preventing misuse and upholding ethical standards. This competition focused on two distinct tracks designed to evaluate and enhance the robustness of LLM security frameworks. Track 1 tasked participants with developing automated methods to probe LLM vulnerabilities by eliciting undesirable responses, effectively testing the limits of existing safety protocols within LLMs. Participants were challenged to devise techniques that could bypass content safeguards across a diverse array of scenarios, from offensive language to misinformation and illegal activities. Through this process, Track 1 aimed to deepen the understanding of LLM vulnerabilities and provide insights for creating more resilient models.
title Global Challenge for Safe and Secure LLMs Track 1
topic Cryptography and Security
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2411.14502