Global Challenge for Safe and Secure LLMs Track 1
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909399637819392 |
|---|---|
| author | Jia, Xiaojun Huang, Yihao Liu, Yang Tan, Peng Yan Yau, Weng Kuan Mak, Mun-Thye Sim, Xin Ming Ng, Wee Siong Ng, See Kiong Liu, Hanqing Zhou, Lifeng Yan, Huanqian Sun, Xiaobing Liu, Wei Wang, Long Qian, Yiming Liu, Yong Yang, Junxiao Zhang, Zhexin Lei, Leqi Chen, Renmiao Lu, Yida Cui, Shiyao Wang, Zizhou Li, Shaohua Wang, Yan Goh, Rick Siow Mong Zhen, Liangli Zhang, Yingjie Zhao, Zhe |
| author_facet | Jia, Xiaojun Huang, Yihao Liu, Yang Tan, Peng Yan Yau, Weng Kuan Mak, Mun-Thye Sim, Xin Ming Ng, Wee Siong Ng, See Kiong Liu, Hanqing Zhou, Lifeng Yan, Huanqian Sun, Xiaobing Liu, Wei Wang, Long Qian, Yiming Liu, Yong Yang, Junxiao Zhang, Zhexin Lei, Leqi Chen, Renmiao Lu, Yida Cui, Shiyao Wang, Zizhou Li, Shaohua Wang, Yan Goh, Rick Siow Mong Zhen, Liangli Zhang, Yingjie Zhao, Zhe |
| contents | This paper introduces the Global Challenge for Safe and Secure Large Language Models (LLMs), a pioneering initiative organized by AI Singapore (AISG) and the CyberSG R&D Programme Office (CRPO) to foster the development of advanced defense mechanisms against automated jailbreaking attacks. With the increasing integration of LLMs in critical sectors such as healthcare, finance, and public administration, ensuring these models are resilient to adversarial attacks is vital for preventing misuse and upholding ethical standards. This competition focused on two distinct tracks designed to evaluate and enhance the robustness of LLM security frameworks. Track 1 tasked participants with developing automated methods to probe LLM vulnerabilities by eliciting undesirable responses, effectively testing the limits of existing safety protocols within LLMs. Participants were challenged to devise techniques that could bypass content safeguards across a diverse array of scenarios, from offensive language to misinformation and illegal activities. Through this process, Track 1 aimed to deepen the understanding of LLM vulnerabilities and provide insights for creating more resilient models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_14502 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Global Challenge for Safe and Secure LLMs Track 1 Jia, Xiaojun Huang, Yihao Liu, Yang Tan, Peng Yan Yau, Weng Kuan Mak, Mun-Thye Sim, Xin Ming Ng, Wee Siong Ng, See Kiong Liu, Hanqing Zhou, Lifeng Yan, Huanqian Sun, Xiaobing Liu, Wei Wang, Long Qian, Yiming Liu, Yong Yang, Junxiao Zhang, Zhexin Lei, Leqi Chen, Renmiao Lu, Yida Cui, Shiyao Wang, Zizhou Li, Shaohua Wang, Yan Goh, Rick Siow Mong Zhen, Liangli Zhang, Yingjie Zhao, Zhe Cryptography and Security Artificial Intelligence Computers and Society This paper introduces the Global Challenge for Safe and Secure Large Language Models (LLMs), a pioneering initiative organized by AI Singapore (AISG) and the CyberSG R&D Programme Office (CRPO) to foster the development of advanced defense mechanisms against automated jailbreaking attacks. With the increasing integration of LLMs in critical sectors such as healthcare, finance, and public administration, ensuring these models are resilient to adversarial attacks is vital for preventing misuse and upholding ethical standards. This competition focused on two distinct tracks designed to evaluate and enhance the robustness of LLM security frameworks. Track 1 tasked participants with developing automated methods to probe LLM vulnerabilities by eliciting undesirable responses, effectively testing the limits of existing safety protocols within LLMs. Participants were challenged to devise techniques that could bypass content safeguards across a diverse array of scenarios, from offensive language to misinformation and illegal activities. Through this process, Track 1 aimed to deepen the understanding of LLM vulnerabilities and provide insights for creating more resilient models. |
| title | Global Challenge for Safe and Secure LLMs Track 1 |
| topic | Cryptography and Security Artificial Intelligence Computers and Society |
| url | https://arxiv.org/abs/2411.14502 |