OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912858499973120 |
|---|---|
| author | Huang, Yu-Shin Just, Peter Yin, Hanyun Narayanan, Krishna Huang, Ruihong Tian, Chao |
| author_facet | Huang, Yu-Shin Just, Peter Yin, Hanyun Narayanan, Krishna Huang, Ruihong Tian, Chao |
| contents | We consider coverless steganography where a Large Language Model (LLM) is used to generate stego-texts in combination with arithmetic coding. An efficient method should embed secret bits in as few language tokens as possible while keeping the stego-text as natural as possible. We show that this problem is equivalent to maximizing the entropy of a replacement probability distribution of the next token generation, subject to a constraint on the divergence between the new distribution and the original one produced by the LLM. A closed-form solution is provided under either the KL divergence or the total variation constraint. Several important practical issues are also tackled: 1) An often-overlooked tokenization mismatch issue is resolved with a simple prompt selection approach, 2) The combination of the optimized distribution and the vocabulary truncation technique is considered, and 3) The incorporation of the proposed approach with existing (potentially non arithmetic coding based) techniques, e.g., the Discop technique. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_04328 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions Huang, Yu-Shin Just, Peter Yin, Hanyun Narayanan, Krishna Huang, Ruihong Tian, Chao Information Theory Artificial Intelligence Computation and Language Cryptography and Security Machine Learning We consider coverless steganography where a Large Language Model (LLM) is used to generate stego-texts in combination with arithmetic coding. An efficient method should embed secret bits in as few language tokens as possible while keeping the stego-text as natural as possible. We show that this problem is equivalent to maximizing the entropy of a replacement probability distribution of the next token generation, subject to a constraint on the divergence between the new distribution and the original one produced by the LLM. A closed-form solution is provided under either the KL divergence or the total variation constraint. Several important practical issues are also tackled: 1) An often-overlooked tokenization mismatch issue is resolved with a simple prompt selection approach, 2) The combination of the optimized distribution and the vocabulary truncation technique is considered, and 3) The incorporation of the proposed approach with existing (potentially non arithmetic coding based) techniques, e.g., the Discop technique. |
| title | OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions |
| topic | Information Theory Artificial Intelligence Computation and Language Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2410.04328 |