Evaluation empirique de la sécurisation et de l'alignement de ChatGPT et Gemini: analyse comparative des vulnérabilités par expérimentations de jailbreaks
Fuente:
arXiv
Salvato in:
| Autore principale: | Nouailles, Rafaël |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects
di: Rahman, Anichur, et al.
Pubblicazione: (2025)
di: Rahman, Anichur, et al.
Pubblicazione: (2025)
AuditGPT: Auditing Smart Contracts with ChatGPT
di: Xia, Shihao, et al.
Pubblicazione: (2024)
di: Xia, Shihao, et al.
Pubblicazione: (2024)
Exploring ChatGPT's Capabilities on Vulnerability Management
di: Liu, Peiyu, et al.
Pubblicazione: (2023)
di: Liu, Peiyu, et al.
Pubblicazione: (2023)
Exfiltration of personal information from ChatGPT via prompt injection
di: Schwartzman, Gregory
Pubblicazione: (2024)
di: Schwartzman, Gregory
Pubblicazione: (2024)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
di: Iqbal, Umar, et al.
Pubblicazione: (2023)
di: Iqbal, Umar, et al.
Pubblicazione: (2023)
On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic Writing
di: Liu, Zeyan, et al.
Pubblicazione: (2023)
di: Liu, Zeyan, et al.
Pubblicazione: (2023)
From Chatbots to PhishBots? -- Preventing Phishing scams created using ChatGPT, Google Bard and Claude
di: Roy, Sayak Saha, et al.
Pubblicazione: (2023)
di: Roy, Sayak Saha, et al.
Pubblicazione: (2023)
ChatGPT's Potential in Cryptography Misuse Detection: A Comparative Analysis with Static Analysis Tools
di: Firouzi, Ehsan, et al.
Pubblicazione: (2024)
di: Firouzi, Ehsan, et al.
Pubblicazione: (2024)
Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments
di: Mukherjee, Kunal, et al.
Pubblicazione: (2026)
di: Mukherjee, Kunal, et al.
Pubblicazione: (2026)
Detecting Phishing Sites Using ChatGPT
di: Koide, Takashi, et al.
Pubblicazione: (2023)
di: Koide, Takashi, et al.
Pubblicazione: (2023)
How Secure is Code Generated by ChatGPT?
di: Khoury, Raphaël, et al.
Pubblicazione: (2023)
di: Khoury, Raphaël, et al.
Pubblicazione: (2023)
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
di: Yi, Xin, et al.
Pubblicazione: (2025)
di: Yi, Xin, et al.
Pubblicazione: (2025)
Can ChatGPT Detect DeepFakes? A Study of Using Multimodal Large Language Models for Media Forensics
di: Jia, Shan, et al.
Pubblicazione: (2024)
di: Jia, Shan, et al.
Pubblicazione: (2024)
Security Analysis of ChatGPT: Threats and Privacy Risks
di: Xiang, Yushan, et al.
Pubblicazione: (2025)
di: Xiang, Yushan, et al.
Pubblicazione: (2025)
Digital Forensic Investigation of the ChatGPT Windows Application
di: Kankanamge, Malithi Wanniarachchi, et al.
Pubblicazione: (2025)
di: Kankanamge, Malithi Wanniarachchi, et al.
Pubblicazione: (2025)
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
di: Shavit, Doron
Pubblicazione: (2026)
di: Shavit, Doron
Pubblicazione: (2026)
A Qualitative Study on Using ChatGPT for Software Security: Perception vs. Practicality
di: Kholoosi, M. Mehdi, et al.
Pubblicazione: (2024)
di: Kholoosi, M. Mehdi, et al.
Pubblicazione: (2024)
EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection
di: Liu, Shigang, et al.
Pubblicazione: (2024)
di: Liu, Shigang, et al.
Pubblicazione: (2024)
Evaluation of ChatGPT's Smart Contract Auditing Capabilities Based on Chain of Thought
di: Du, Yuying, et al.
Pubblicazione: (2024)
di: Du, Yuying, et al.
Pubblicazione: (2024)
Are aligned neural networks adversarially aligned?
di: Carlini, Nicholas, et al.
Pubblicazione: (2023)
di: Carlini, Nicholas, et al.
Pubblicazione: (2023)
Exploring Backdoor Vulnerabilities of Chat Models
di: Hao, Yunzhuo, et al.
Pubblicazione: (2024)
di: Hao, Yunzhuo, et al.
Pubblicazione: (2024)
Time to Separate from StackOverflow and Match with ChatGPT for Encryption
di: Firouzi, Ehsan, et al.
Pubblicazione: (2024)
di: Firouzi, Ehsan, et al.
Pubblicazione: (2024)
Can ChatGPT Perform Image Splicing Detection? A Preliminary Study
di: Nath, Souradip
Pubblicazione: (2025)
di: Nath, Souradip
Pubblicazione: (2025)
Just another copy and paste? Comparing the security vulnerabilities of ChatGPT generated code and StackOverflow answers
di: Hamer, Sivana, et al.
Pubblicazione: (2024)
di: Hamer, Sivana, et al.
Pubblicazione: (2024)
ChatGPT: Excellent Paper! Accept It. Editor: Imposter Found! Review Rejected
di: Gharami, Kanchon, et al.
Pubblicazione: (2025)
di: Gharami, Kanchon, et al.
Pubblicazione: (2025)
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
di: Wang, Boxin, et al.
Pubblicazione: (2023)
di: Wang, Boxin, et al.
Pubblicazione: (2023)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
di: Wang, Chao, et al.
Pubblicazione: (2026)
di: Wang, Chao, et al.
Pubblicazione: (2026)
WildCode: An Empirical Analysis of Code Generated by ChatGPT
di: Khanmohammadi, Kobra, et al.
Pubblicazione: (2025)
di: Khanmohammadi, Kobra, et al.
Pubblicazione: (2025)
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
di: Ramesh, Govind, et al.
Pubblicazione: (2024)
di: Ramesh, Govind, et al.
Pubblicazione: (2024)
ChatGPT, is this real? The influence of generative AI on writing style in top-tier cybersecurity papers
di: Vansteenhuyse, Daan
Pubblicazione: (2026)
di: Vansteenhuyse, Daan
Pubblicazione: (2026)
AbuseGPT: Abuse of Generative AI ChatBots to Create Smishing Campaigns
di: Shibli, Ashfak Md, et al.
Pubblicazione: (2024)
di: Shibli, Ashfak Md, et al.
Pubblicazione: (2024)
Enhancing Android Malware Detection: The Influence of ChatGPT on Decision-centric Task
di: Li, Yao, et al.
Pubblicazione: (2024)
di: Li, Yao, et al.
Pubblicazione: (2024)
ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications
di: Zaboli, Aydin, et al.
Pubblicazione: (2023)
di: Zaboli, Aydin, et al.
Pubblicazione: (2023)
From static to adaptive: immune memory-based jailbreak detection for large language models
di: Leng, Jun, et al.
Pubblicazione: (2025)
di: Leng, Jun, et al.
Pubblicazione: (2025)
Low-Resource Languages Jailbreak GPT-4
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2023)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2023)
Using Hallucinations to Bypass GPT4's Filter
di: Lemkin, Benjamin
Pubblicazione: (2024)
di: Lemkin, Benjamin
Pubblicazione: (2024)
Generative AI like ChatGPT in Blockchain Federated Learning: use cases, opportunities and future
di: Puppala, Sai, et al.
Pubblicazione: (2024)
di: Puppala, Sai, et al.
Pubblicazione: (2024)
Breaking the Prompt Wall (I): A Real-World Case Study of Attacking ChatGPT via Lightweight Prompt Injection
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
AVISE: Framework for Evaluating the Security of AI Systems
di: Lempinen, Mikko, et al.
Pubblicazione: (2026)
di: Lempinen, Mikko, et al.
Pubblicazione: (2026)
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
di: Chen, Xiuyuan, et al.
Pubblicazione: (2025)
di: Chen, Xiuyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects
di: Rahman, Anichur, et al.
Pubblicazione: (2025) -
AuditGPT: Auditing Smart Contracts with ChatGPT
di: Xia, Shihao, et al.
Pubblicazione: (2024) -
Exploring ChatGPT's Capabilities on Vulnerability Management
di: Liu, Peiyu, et al.
Pubblicazione: (2023) -
Exfiltration of personal information from ChatGPT via prompt injection
di: Schwartzman, Gregory
Pubblicazione: (2024) -
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
di: Iqbal, Umar, et al.
Pubblicazione: (2023)