Early Signs of Steganographic Capabilities in Frontier LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zolkowski, Artur, Nishimura-Gasparian, Kei, McCarthy, Robert, Zimmermann, Roland S., Lindner, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
The Steganographic Potentials of Language Models
von: Karpov, Artem, et al.
Veröffentlicht: (2025)
von: Karpov, Artem, et al.
Veröffentlicht: (2025)
Towards Understanding Specification Gaming in Reasoning Models
von: Nishimura-Gasparian, Kei, et al.
Veröffentlicht: (2026)
von: Nishimura-Gasparian, Kei, et al.
Veröffentlicht: (2026)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
Jailbroken Frontier Models Retain Their Capabilities
von: Zhu, Daniel, et al.
Veröffentlicht: (2026)
von: Zhu, Daniel, et al.
Veröffentlicht: (2026)
An Independent Safety Evaluation of Kimi K2.5
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2026)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2026)
NEST: Nascent Encoded Steganographic Thoughts
von: Karpov, Artem
Veröffentlicht: (2026)
von: Karpov, Artem
Veröffentlicht: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
There Are No Silly Questions: Evaluation of Offline LLM Capabilities from a Turkish Perspective
von: Yilmaz, Edibe, et al.
Veröffentlicht: (2026)
von: Yilmaz, Edibe, et al.
Veröffentlicht: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Gandalf the Red: Adaptive Security for LLMs
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
Jailbreaking LLMs via Calibration
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
von: Tong, Xin, et al.
Veröffentlicht: (2025)
von: Tong, Xin, et al.
Veröffentlicht: (2025)
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
von: Marchand, Rahul, et al.
Veröffentlicht: (2026)
von: Marchand, Rahul, et al.
Veröffentlicht: (2026)
Tool Preferences in Agentic LLMs are Unreliable
von: Faghih, Kazem, et al.
Veröffentlicht: (2025)
von: Faghih, Kazem, et al.
Veröffentlicht: (2025)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
von: Cögendez, Derya, et al.
Veröffentlicht: (2026)
von: Cögendez, Derya, et al.
Veröffentlicht: (2026)
Shh, don't say that! Domain Certification in LLMs
von: Emde, Cornelius, et al.
Veröffentlicht: (2025)
von: Emde, Cornelius, et al.
Veröffentlicht: (2025)
SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs
von: Pan, Chao, et al.
Veröffentlicht: (2026)
von: Pan, Chao, et al.
Veröffentlicht: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
Jailbreaking Frontier Foundation Models Through Intention Deception
von: Wang, Xinhe, et al.
Veröffentlicht: (2026)
von: Wang, Xinhe, et al.
Veröffentlicht: (2026)
Internal Safety Collapse in Frontier Large Language Models
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
Beyond Slow Signs in High-fidelity Model Extraction
von: Foerster, Hanna, et al.
Veröffentlicht: (2024)
von: Foerster, Hanna, et al.
Veröffentlicht: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
von: Vega, Jason, et al.
Veröffentlicht: (2023)
von: Vega, Jason, et al.
Veröffentlicht: (2023)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protection
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
BaxBench: Can LLMs Generate Correct and Secure Backends?
von: Vero, Mark, et al.
Veröffentlicht: (2025)
von: Vero, Mark, et al.
Veröffentlicht: (2025)
LLMs can hide text in other text of the same length
von: Norelli, Antonio, et al.
Veröffentlicht: (2025)
von: Norelli, Antonio, et al.
Veröffentlicht: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
von: Kan, Chun Yan Ryan, et al.
Veröffentlicht: (2026)
von: Kan, Chun Yan Ryan, et al.
Veröffentlicht: (2026)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025) -
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
von: Mathew, Yohan, et al.
Veröffentlicht: (2024) -
The Steganographic Potentials of Language Models
von: Karpov, Artem, et al.
Veröffentlicht: (2025) -
Towards Understanding Specification Gaming in Reasoning Models
von: Nishimura-Gasparian, Kei, et al.
Veröffentlicht: (2026) -
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)