Relatively-Secure LLM-Based Steganography via Constrained Markov Decision Processes

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Yu-Shin, Tian, Chao, Narayanan, Krishna, Zheng, Lizhong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910812110585856
author Huang, Yu-Shin
Tian, Chao
Narayanan, Krishna
Zheng, Lizhong
author_facet Huang, Yu-Shin
Tian, Chao
Narayanan, Krishna
Zheng, Lizhong
contents Linguistic steganography aims to conceal information within natural language text without being detected. An effective steganography approach should encode the secret message into a minimal number of language tokens while preserving the natural appearance and fluidity of the stego-texts. We present a new framework to enhance the embedding efficiency of stego-texts generated by modifying the output of a large language model (LLM). The novelty of our approach is in abstracting the sequential steganographic embedding process as a Constrained Markov Decision Process (CMDP), which takes into consideration the long-term dependencies instead of merely the immediate effects. We constrain the solution space such that the discounted accumulative total variation divergence between the selected probability distribution and the original distribution given by the LLM is below a threshold. To find the optimal policy, we first show that the functional optimization problem can be simplified to a convex optimization problem with a finite number of variables. A closed-form solution for the optimal policy is then presented to this equivalent problem. It is remarkable that the optimal policy is deterministic and resembles water-filling in some cases. The solution suggests that usually adjusting the probability distribution for the state that has the least random transition probability should be prioritized, but the choice should be made by taking into account the transition probabilities at all states instead of only the current state.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Relatively-Secure LLM-Based Steganography via Constrained Markov Decision Processes
Huang, Yu-Shin
Tian, Chao
Narayanan, Krishna
Zheng, Lizhong
Information Theory
Linguistic steganography aims to conceal information within natural language text without being detected. An effective steganography approach should encode the secret message into a minimal number of language tokens while preserving the natural appearance and fluidity of the stego-texts. We present a new framework to enhance the embedding efficiency of stego-texts generated by modifying the output of a large language model (LLM). The novelty of our approach is in abstracting the sequential steganographic embedding process as a Constrained Markov Decision Process (CMDP), which takes into consideration the long-term dependencies instead of merely the immediate effects. We constrain the solution space such that the discounted accumulative total variation divergence between the selected probability distribution and the original distribution given by the LLM is below a threshold. To find the optimal policy, we first show that the functional optimization problem can be simplified to a convex optimization problem with a finite number of variables. A closed-form solution for the optimal policy is then presented to this equivalent problem. It is remarkable that the optimal policy is deterministic and resembles water-filling in some cases. The solution suggests that usually adjusting the probability distribution for the state that has the least random transition probability should be prioritized, but the choice should be made by taking into account the transition probabilities at all states instead of only the current state.
title Relatively-Secure LLM-Based Steganography via Constrained Markov Decision Processes
topic Information Theory
url https://arxiv.org/abs/2502.01827