BadGPT-4o: stripping safety finetuning from GPT models
Fuente:
arXiv
Saved in:
| Main Authors: | Krupkina, Ekaterina, Volkov, Dmitrii |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Badllama 3: removing safety finetuning from Llama 3 in minutes
by: Volkov, Dmitrii
Published: (2024)
by: Volkov, Dmitrii
Published: (2024)
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024)
by: Shen, Xinyue, et al.
Published: (2024)
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025)
by: Reworr, et al.
Published: (2025)
An Efficient Private GPT Never Autoregressively Decodes
by: Li, Zhengyi, et al.
Published: (2025)
by: Li, Zhengyi, et al.
Published: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024)
by: Chen, Shuo, et al.
Published: (2024)
Initial Exploration of Zero-Shot Privacy Utility Tradeoffs in Tabular Data Using GPT-4
by: Mandal, Bishwas, et al.
Published: (2024)
by: Mandal, Bishwas, et al.
Published: (2024)
A Novel GPT-Based Framework for Anomaly Detection in System Logs
by: Zhang, Zeng, et al.
Published: (2025)
by: Zhang, Zeng, et al.
Published: (2025)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
IstGPT: LLM-based Anomaly Detection for Spatial-Temporal Graph in Industrial Systems
by: Zhang, Yuchen, et al.
Published: (2026)
by: Zhang, Yuchen, et al.
Published: (2026)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
by: Wu, Yuanwei, et al.
Published: (2023)
by: Wu, Yuanwei, et al.
Published: (2023)
Using Hallucinations to Bypass GPT4's Filter
by: Lemkin, Benjamin
Published: (2024)
by: Lemkin, Benjamin
Published: (2024)
Low-Resource Languages Jailbreak GPT-4
by: Yong, Zheng-Xin, et al.
Published: (2023)
by: Yong, Zheng-Xin, et al.
Published: (2023)
Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense: A 2022 Study of GPT-3 and Contemporary Models
by: Sandoval, Gustavo, et al.
Published: (2025)
by: Sandoval, Gustavo, et al.
Published: (2025)
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024)
by: Zhang, Jinghuai, et al.
Published: (2024)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
by: Xiang, Zhen, et al.
Published: (2024)
by: Xiang, Zhen, et al.
Published: (2024)
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
by: Agarwal, Krishiv, et al.
Published: (2026)
by: Agarwal, Krishiv, et al.
Published: (2026)
BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
by: Lu, Bingguang, et al.
Published: (2025)
by: Lu, Bingguang, et al.
Published: (2025)
BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models
by: Wang, Jiayao, et al.
Published: (2026)
by: Wang, Jiayao, et al.
Published: (2026)
BadGD: A unified data-centric framework to identify gradient descent vulnerabilities
by: Wang, Chi-Hua, et al.
Published: (2024)
by: Wang, Chi-Hua, et al.
Published: (2024)
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
by: Zhang, Mingxuan, et al.
Published: (2025)
by: Zhang, Mingxuan, et al.
Published: (2025)
On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic Writing
by: Liu, Zeyan, et al.
Published: (2023)
by: Liu, Zeyan, et al.
Published: (2023)
Software Vulnerability Prediction in Low-Resource Languages: An Empirical Study of CodeBERT and ChatGPT
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023)
by: Pelrine, Kellin, et al.
Published: (2023)
Naturally Private Recommendations with Determinantal Point Processes
by: Fitzsimons, Jack, et al.
Published: (2024)
by: Fitzsimons, Jack, et al.
Published: (2024)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
Federated Learning Clients Clustering with Adaptation to Data Drifts
by: Li, Minghao, et al.
Published: (2024)
by: Li, Minghao, et al.
Published: (2024)
HW-V2W-Map: Hardware Vulnerability to Weakness Mapping Framework for Root Cause Analysis with GPT-assisted Mitigation Suggestion
by: Lin, Yu-Zheng, et al.
Published: (2023)
by: Lin, Yu-Zheng, et al.
Published: (2023)
Virtual camera detection: Catching video injection attacks in remote biometric systems
by: Kurmankhojayev, Daniyar, et al.
Published: (2025)
by: Kurmankhojayev, Daniyar, et al.
Published: (2025)
BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
Evaluating AI cyber capabilities with crowdsourced elicitation
by: Petrov, Artem, et al.
Published: (2025)
by: Petrov, Artem, et al.
Published: (2025)
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
by: Usynin, Dmitrii, et al.
Published: (2023)
by: Usynin, Dmitrii, et al.
Published: (2023)
BadSampler: Harnessing the Power of Catastrophic Forgetting to Poison Byzantine-robust Federated Learning
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
A Cognac Shot To Forget Bad Memories: Corrective Unlearning for Graph Neural Networks
by: Kolipaka, Varshita, et al.
Published: (2024)
by: Kolipaka, Varshita, et al.
Published: (2024)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
Similar Items
-
Badllama 3: removing safety finetuning from Llama 3 in minutes
by: Volkov, Dmitrii
Published: (2024) -
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024) -
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
by: Reworr, et al.
Published: (2025) -
An Efficient Private GPT Never Autoregressively Decodes
by: Li, Zhengyi, et al.
Published: (2025) -
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024)