Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4
Fuente:
arXiv
Salvato in:
| Autori principali: | Polyakov, Alex, Kuznetsov, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Local Frames: Exploiting Inherited Origins to Bypass Content Blockers
di: Ukani, Alisha, et al.
Pubblicazione: (2025)
di: Ukani, Alisha, et al.
Pubblicazione: (2025)
Few-Shot Learning-Based Cyber Incident Detection with Augmented Context Intelligence
di: Zuo, Fei, et al.
Pubblicazione: (2025)
di: Zuo, Fei, et al.
Pubblicazione: (2025)
Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning
di: Komaty, Alain, et al.
Pubblicazione: (2025)
di: Komaty, Alain, et al.
Pubblicazione: (2025)
Generalized Adversarial Code-Suggestions: Exploiting Contexts of LLM-based Code-Completion
di: Rubel, Karl, et al.
Pubblicazione: (2024)
di: Rubel, Karl, et al.
Pubblicazione: (2024)
WAFFLED: Exploiting Parsing Discrepancies to Bypass Web Application Firewalls
di: Akhavani, Seyed Ali, et al.
Pubblicazione: (2025)
di: Akhavani, Seyed Ali, et al.
Pubblicazione: (2025)
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
di: Yuan, Shuai, et al.
Pubblicazione: (2025)
di: Yuan, Shuai, et al.
Pubblicazione: (2025)
FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
di: Kuznetsov, Daniel, et al.
Pubblicazione: (2026)
di: Kuznetsov, Daniel, et al.
Pubblicazione: (2026)
Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation
di: Tang, Xinyu, et al.
Pubblicazione: (2023)
di: Tang, Xinyu, et al.
Pubblicazione: (2023)
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
di: Li, Junchen, et al.
Pubblicazione: (2026)
di: Li, Junchen, et al.
Pubblicazione: (2026)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
di: Zaree, Pedram, et al.
Pubblicazione: (2025)
di: Zaree, Pedram, et al.
Pubblicazione: (2025)
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
di: Xu, Zhenhao, et al.
Pubblicazione: (2026)
di: Xu, Zhenhao, et al.
Pubblicazione: (2026)
Involuntary Jailbreak: On Self-Prompting Attacks
di: Guo, Yangyang, et al.
Pubblicazione: (2025)
di: Guo, Yangyang, et al.
Pubblicazione: (2025)
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
di: Halloran, John
Pubblicazione: (2025)
di: Halloran, John
Pubblicazione: (2025)
Physical-Layer Signal Injection Attacks on EV Charging Ports: Bypassing Authentication via Electrical-Level Exploits
di: Shi, Hetian, et al.
Pubblicazione: (2025)
di: Shi, Hetian, et al.
Pubblicazione: (2025)
Using Hallucinations to Bypass GPT4's Filter
di: Lemkin, Benjamin
Pubblicazione: (2024)
di: Lemkin, Benjamin
Pubblicazione: (2024)
Distributed Intrusion Detection in Dynamic Networks of UAVs using Few-Shot Federated Learning
di: Ceviz, Ozlem, et al.
Pubblicazione: (2025)
di: Ceviz, Ozlem, et al.
Pubblicazione: (2025)
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
di: He, Sinan, et al.
Pubblicazione: (2025)
di: He, Sinan, et al.
Pubblicazione: (2025)
In-Context Unlearning: Language Models as Few Shot Unlearners
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
Model X-Ray: Detection of Hidden Malware in AI Model Weights using Few Shot Learning
di: Gilkarov, Daniel, et al.
Pubblicazione: (2024)
di: Gilkarov, Daniel, et al.
Pubblicazione: (2024)
FCert: Certifiably Robust Few-Shot Classification in the Era of Foundation Models
di: Wang, Yanting, et al.
Pubblicazione: (2024)
di: Wang, Yanting, et al.
Pubblicazione: (2024)
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
di: Shi, Yunfan
Pubblicazione: (2024)
di: Shi, Yunfan
Pubblicazione: (2024)
VIPER-MCP: Detecting and Exploiting Taint-Style Vulnerabilities in Model Context Protocol Servers
di: Sun, Pengyu, et al.
Pubblicazione: (2026)
di: Sun, Pengyu, et al.
Pubblicazione: (2026)
SleepWalk: Exploiting Context Switching and Residual Power for Physical Side-Channel Attacks
di: Sanjaya, Sahan, et al.
Pubblicazione: (2025)
di: Sanjaya, Sahan, et al.
Pubblicazione: (2025)
Agent Safety Alignment via Reinforcement Learning
di: Sha, Zeyang, et al.
Pubblicazione: (2025)
di: Sha, Zeyang, et al.
Pubblicazione: (2025)
MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
di: Radosevich, Brandon, et al.
Pubblicazione: (2025)
di: Radosevich, Brandon, et al.
Pubblicazione: (2025)
Towards Novel Malicious Packet Recognition: A Few-Shot Learning Approach
di: Stein, Kyle, et al.
Pubblicazione: (2024)
di: Stein, Kyle, et al.
Pubblicazione: (2024)
Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing
di: Holtzman, Ari, et al.
Pubblicazione: (2026)
di: Holtzman, Ari, et al.
Pubblicazione: (2026)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
A Classification-by-Retrieval Framework for Few-Shot Anomaly Detection to Detect API Injection Attacks
di: Aharon, Udi, et al.
Pubblicazione: (2024)
di: Aharon, Udi, et al.
Pubblicazione: (2024)
Strengthening Network Intrusion Detection in IoT Environments with Self-Supervised Learning and Few Shot Learning
di: Atitallah, Safa Ben, et al.
Pubblicazione: (2024)
di: Atitallah, Safa Ben, et al.
Pubblicazione: (2024)
Few Edges Are Enough: Few-Shot Network Attack Detection with Graph Neural Networks
di: Bilot, Tristan, et al.
Pubblicazione: (2025)
di: Bilot, Tristan, et al.
Pubblicazione: (2025)
Efficient Denial of Service Attack Detection in IoT using Kolmogorov-Arnold Networks
di: Kuznetsov, Oleksandr
Pubblicazione: (2025)
di: Kuznetsov, Oleksandr
Pubblicazione: (2025)
MalMixer: Few-Shot Malware Classification with Retrieval-Augmented Semi-Supervised Learning
di: Li, Jiliang, et al.
Pubblicazione: (2024)
di: Li, Jiliang, et al.
Pubblicazione: (2024)
TREC: APT Tactic / Technique Recognition via Few-Shot Provenance Subgraph Learning
di: Lv, Mingqi, et al.
Pubblicazione: (2024)
di: Lv, Mingqi, et al.
Pubblicazione: (2024)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
di: Huang, Chao, et al.
Pubblicazione: (2025)
di: Huang, Chao, et al.
Pubblicazione: (2025)
Initial Exploration of Zero-Shot Privacy Utility Tradeoffs in Tabular Data Using GPT-4
di: Mandal, Bishwas, et al.
Pubblicazione: (2024)
di: Mandal, Bishwas, et al.
Pubblicazione: (2024)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
di: Wang, Chao, et al.
Pubblicazione: (2026)
di: Wang, Chao, et al.
Pubblicazione: (2026)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIs
di: Zhang, Jianqing, et al.
Pubblicazione: (2025)
di: Zhang, Jianqing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Local Frames: Exploiting Inherited Origins to Bypass Content Blockers
di: Ukani, Alisha, et al.
Pubblicazione: (2025) -
Few-Shot Learning-Based Cyber Incident Detection with Augmented Context Intelligence
di: Zuo, Fei, et al.
Pubblicazione: (2025) -
Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning
di: Komaty, Alain, et al.
Pubblicazione: (2025) -
Generalized Adversarial Code-Suggestions: Exploiting Contexts of LLM-based Code-Completion
di: Rubel, Karl, et al.
Pubblicazione: (2024) -
WAFFLED: Exploiting Parsing Discrepancies to Bypass Web Application Firewalls
di: Akhavani, Seyed Ali, et al.
Pubblicazione: (2025)