OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yu-Shin, Just, Peter, Yin, Hanyun, Narayanan, Krishna, Huang, Ruihong, Tian, Chao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912858499973120
author Huang, Yu-Shin
Just, Peter
Yin, Hanyun
Narayanan, Krishna
Huang, Ruihong
Tian, Chao
author_facet Huang, Yu-Shin
Just, Peter
Yin, Hanyun
Narayanan, Krishna
Huang, Ruihong
Tian, Chao
contents We consider coverless steganography where a Large Language Model (LLM) is used to generate stego-texts in combination with arithmetic coding. An efficient method should embed secret bits in as few language tokens as possible while keeping the stego-text as natural as possible. We show that this problem is equivalent to maximizing the entropy of a replacement probability distribution of the next token generation, subject to a constraint on the divergence between the new distribution and the original one produced by the LLM. A closed-form solution is provided under either the KL divergence or the total variation constraint. Several important practical issues are also tackled: 1) An often-overlooked tokenization mismatch issue is resolved with a simple prompt selection approach, 2) The combination of the optimized distribution and the vocabulary truncation technique is considered, and 3) The incorporation of the proposed approach with existing (potentially non arithmetic coding based) techniques, e.g., the Discop technique.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04328
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions
Huang, Yu-Shin
Just, Peter
Yin, Hanyun
Narayanan, Krishna
Huang, Ruihong
Tian, Chao
Information Theory
Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
We consider coverless steganography where a Large Language Model (LLM) is used to generate stego-texts in combination with arithmetic coding. An efficient method should embed secret bits in as few language tokens as possible while keeping the stego-text as natural as possible. We show that this problem is equivalent to maximizing the entropy of a replacement probability distribution of the next token generation, subject to a constraint on the divergence between the new distribution and the original one produced by the LLM. A closed-form solution is provided under either the KL divergence or the total variation constraint. Several important practical issues are also tackled: 1) An often-overlooked tokenization mismatch issue is resolved with a simple prompt selection approach, 2) The combination of the optimized distribution and the vocabulary truncation technique is considered, and 3) The incorporation of the proposed approach with existing (potentially non arithmetic coding based) techniques, e.g., the Discop technique.
title OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions
topic Information Theory
Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2410.04328