One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Tan, Yixin, Yu, Zhe, Sakuma, Jun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
por: Junhao, Wei, et al.
Publicado: (2025)
por: Junhao, Wei, et al.
Publicado: (2025)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
por: Qi, Senmao, et al.
Publicado: (2025)
por: Qi, Senmao, et al.
Publicado: (2025)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
por: Nihal, Ragib Amin, et al.
Publicado: (2025)
por: Nihal, Ragib Amin, et al.
Publicado: (2025)
New Wide-Net-Casting Jailbreak Attacks Risk Large Models
por: Xiang, Qiuchi, et al.
Publicado: (2026)
por: Xiang, Qiuchi, et al.
Publicado: (2026)
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
por: Yan, Yu, et al.
Publicado: (2025)
por: Yan, Yu, et al.
Publicado: (2025)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
por: Lu, Yu-An, et al.
Publicado: (2026)
por: Lu, Yu-An, et al.
Publicado: (2026)
FlipAttack: Jailbreak LLMs via Flipping
por: Liu, Yue, et al.
Publicado: (2024)
por: Liu, Yue, et al.
Publicado: (2024)
Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
por: Li, Yakai, et al.
Publicado: (2025)
por: Li, Yakai, et al.
Publicado: (2025)
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
por: Zhao, Wei, et al.
Publicado: (2024)
por: Zhao, Wei, et al.
Publicado: (2024)
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
por: Xiong, Chen, et al.
Publicado: (2026)
por: Xiong, Chen, et al.
Publicado: (2026)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
por: Galinkin, Erick, et al.
Publicado: (2024)
por: Galinkin, Erick, et al.
Publicado: (2024)
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
por: Ahn, Yelim, et al.
Publicado: (2025)
por: Ahn, Yelim, et al.
Publicado: (2025)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
por: Zhang, Zheng, et al.
Publicado: (2025)
por: Zhang, Zheng, et al.
Publicado: (2025)
LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
por: Jaiswal, Piyush, et al.
Publicado: (2026)
por: Jaiswal, Piyush, et al.
Publicado: (2026)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
por: Lin, Zheng, et al.
Publicado: (2026)
por: Lin, Zheng, et al.
Publicado: (2026)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
por: Chen, Kejia, et al.
Publicado: (2026)
por: Chen, Kejia, et al.
Publicado: (2026)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
por: Sel, Bilgehan, et al.
Publicado: (2026)
por: Sel, Bilgehan, et al.
Publicado: (2026)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
por: Teng, Ma, et al.
Publicado: (2024)
por: Teng, Ma, et al.
Publicado: (2024)
Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
por: Wong, Ryan, et al.
Publicado: (2025)
por: Wong, Ryan, et al.
Publicado: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
por: Shang, Zhengchun, et al.
Publicado: (2025)
por: Shang, Zhengchun, et al.
Publicado: (2025)
Alphabet Index Mapping: Jailbreaking LLMs through Semantic Dissimilarity
por: Husain, Bilal Saleh
Publicado: (2025)
por: Husain, Bilal Saleh
Publicado: (2025)
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
por: Cui, Tiehan, et al.
Publicado: (2025)
por: Cui, Tiehan, et al.
Publicado: (2025)
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
por: Gong, Xueluan, et al.
Publicado: (2024)
por: Gong, Xueluan, et al.
Publicado: (2024)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
por: Yoon, Sangyeon, et al.
Publicado: (2026)
por: Yoon, Sangyeon, et al.
Publicado: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
por: Guo, Weiyang, et al.
Publicado: (2026)
por: Guo, Weiyang, et al.
Publicado: (2026)
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
por: Liu, Tong, et al.
Publicado: (2024)
por: Liu, Tong, et al.
Publicado: (2024)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
por: Nikolić, Kristina, et al.
Publicado: (2025)
por: Nikolić, Kristina, et al.
Publicado: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
por: Liu, Xiaoqun, et al.
Publicado: (2024)
por: Liu, Xiaoqun, et al.
Publicado: (2024)
Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs
por: Ling, Zijian, et al.
Publicado: (2026)
por: Ling, Zijian, et al.
Publicado: (2026)
CompLeak: Deep Learning Model Compression Exacerbates Privacy Leakage
por: Li, Na, et al.
Publicado: (2025)
por: Li, Na, et al.
Publicado: (2025)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
por: Zhang, Yingjie, et al.
Publicado: (2025)
por: Zhang, Yingjie, et al.
Publicado: (2025)
TMI! Finetuned Models Leak Private Information from their Pretraining Data
por: Abascal, John, et al.
Publicado: (2023)
por: Abascal, John, et al.
Publicado: (2023)
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
por: Sun, Zhen, et al.
Publicado: (2025)
por: Sun, Zhen, et al.
Publicado: (2025)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
por: Xu, Wenzhuo, et al.
Publicado: (2026)
por: Xu, Wenzhuo, et al.
Publicado: (2026)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
por: Chen, Yu, et al.
Publicado: (2026)
por: Chen, Yu, et al.
Publicado: (2026)
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
por: Chen, Meifang, et al.
Publicado: (2026)
por: Chen, Meifang, et al.
Publicado: (2026)
Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
por: Marinelli, Ryan, et al.
Publicado: (2025)
por: Marinelli, Ryan, et al.
Publicado: (2025)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
por: Li, Xiaohu, et al.
Publicado: (2025)
por: Li, Xiaohu, et al.
Publicado: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
por: Tong, Haibo, et al.
Publicado: (2025)
por: Tong, Haibo, et al.
Publicado: (2025)
Ejemplares similares
-
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
por: Junhao, Wei, et al.
Publicado: (2025) -
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
por: Qi, Senmao, et al.
Publicado: (2025) -
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
por: Nihal, Ragib Amin, et al.
Publicado: (2025) -
New Wide-Net-Casting Jailbreak Attacks Risk Large Models
por: Xiang, Qiuchi, et al.
Publicado: (2026) -
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
por: Yan, Yu, et al.
Publicado: (2025)