LLMs can hide text in other text of the same length
Fuente:
arXiv
Salvato in:
| Autori principali: | Norelli, Antonio, Bronstein, Michael |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
di: Betley, Jan, et al.
Pubblicazione: (2025)
di: Betley, Jan, et al.
Pubblicazione: (2025)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
di: Lin, Shi, et al.
Pubblicazione: (2024)
di: Lin, Shi, et al.
Pubblicazione: (2024)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
di: Dubiński, Jan, et al.
Pubblicazione: (2026)
di: Dubiński, Jan, et al.
Pubblicazione: (2026)
De-identification of clinical free text using natural language processing: A systematic review of current approaches
di: Kovačević, Aleksandar, et al.
Pubblicazione: (2023)
di: Kovačević, Aleksandar, et al.
Pubblicazione: (2023)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
Private prediction for large-scale synthetic text generation
di: Amin, Kareem, et al.
Pubblicazione: (2024)
di: Amin, Kareem, et al.
Pubblicazione: (2024)
Jailbreaking LLMs via Calibration
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
Tool Preferences in Agentic LLMs are Unreliable
di: Faghih, Kazem, et al.
Pubblicazione: (2025)
di: Faghih, Kazem, et al.
Pubblicazione: (2025)
Gandalf the Red: Adaptive Security for LLMs
di: Pfister, Niklas, et al.
Pubblicazione: (2025)
di: Pfister, Niklas, et al.
Pubblicazione: (2025)
Shh, don't say that! Domain Certification in LLMs
di: Emde, Cornelius, et al.
Pubblicazione: (2025)
di: Emde, Cornelius, et al.
Pubblicazione: (2025)
Early Signs of Steganographic Capabilities in Frontier LLMs
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
di: Paulus, Anselm, et al.
Pubblicazione: (2024)
di: Paulus, Anselm, et al.
Pubblicazione: (2024)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
di: Betley, Jan, et al.
Pubblicazione: (2025)
di: Betley, Jan, et al.
Pubblicazione: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
di: Xu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Xu, Xiaoyu, et al.
Pubblicazione: (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
di: Betley, Jan, et al.
Pubblicazione: (2025)
di: Betley, Jan, et al.
Pubblicazione: (2025)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
di: Vega, Jason, et al.
Pubblicazione: (2023)
di: Vega, Jason, et al.
Pubblicazione: (2023)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protection
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
Learning to Diagnose Privately: DP-Powered LLMs for Radiology Report Classification
di: Bhattacharjee, Payel, et al.
Pubblicazione: (2025)
di: Bhattacharjee, Payel, et al.
Pubblicazione: (2025)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
di: Golchin, Shahriar, et al.
Pubblicazione: (2023)
di: Golchin, Shahriar, et al.
Pubblicazione: (2023)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
di: Kan, Chun Yan Ryan, et al.
Pubblicazione: (2026)
di: Kan, Chun Yan Ryan, et al.
Pubblicazione: (2026)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
di: Poppi, Samuele, et al.
Pubblicazione: (2024)
di: Poppi, Samuele, et al.
Pubblicazione: (2024)
Teach LLMs to Phish: Stealing Private Information from Language Models
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
di: Wang, Kai, et al.
Pubblicazione: (2025)
di: Wang, Kai, et al.
Pubblicazione: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
di: Liang, Buyun, et al.
Pubblicazione: (2025)
di: Liang, Buyun, et al.
Pubblicazione: (2025)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
di: Guo, Chuan, et al.
Pubblicazione: (2026)
di: Guo, Chuan, et al.
Pubblicazione: (2026)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
di: Hasan, Adib, et al.
Pubblicazione: (2024)
di: Hasan, Adib, et al.
Pubblicazione: (2024)
How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
di: Mostafa, Ahmed, et al.
Pubblicazione: (2025)
di: Mostafa, Ahmed, et al.
Pubblicazione: (2025)
FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
di: Sarwar, Nobin, et al.
Pubblicazione: (2025)
di: Sarwar, Nobin, et al.
Pubblicazione: (2025)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
di: Frikha, Ahmed, et al.
Pubblicazione: (2024)
di: Frikha, Ahmed, et al.
Pubblicazione: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
di: Alhazbi, Saeif, et al.
Pubblicazione: (2025)
di: Alhazbi, Saeif, et al.
Pubblicazione: (2025)
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
di: Salahuddin, Salahuddin, et al.
Pubblicazione: (2025)
di: Salahuddin, Salahuddin, et al.
Pubblicazione: (2025)
Rethinking How to Evaluate Language Model Jailbreak
di: Cai, Hongyu, et al.
Pubblicazione: (2024)
di: Cai, Hongyu, et al.
Pubblicazione: (2024)
In-Context Representation Hijacking
di: Yona, Itay, et al.
Pubblicazione: (2025)
di: Yona, Itay, et al.
Pubblicazione: (2025)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
di: Betley, Jan, et al.
Pubblicazione: (2025) -
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
di: Lin, Shi, et al.
Pubblicazione: (2024) -
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
di: Dubiński, Jan, et al.
Pubblicazione: (2026) -
De-identification of clinical free text using natural language processing: A systematic review of current approaches
di: Kovačević, Aleksandar, et al.
Pubblicazione: (2023) -
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)