A high-capacity linguistic steganography based on entropy-driven rank-token mapping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Jun, Zhang, Weiming, Yu, Nenghai, Chen, Kejiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917045086453760
author Jiang, Jun
Zhang, Weiming
Yu, Nenghai
Chen, Kejiang
author_facet Jiang, Jun
Zhang, Weiming
Yu, Nenghai
Chen, Kejiang
contents Linguistic steganography enables covert communication through embedding secret messages into innocuous texts; however, current methods face critical limitations in payload capacity and security. Traditional modification-based methods introduce detectable anomalies, while retrieval-based strategies suffer from low embedding capacity. Modern generative steganography leverages language models to generate natural stego text but struggles with limited entropy in token predictions, further constraining capacity. To address these issues, we propose an entropy-driven framework called RTMStega that integrates rank-based adaptive coding and context-aware decompression with normalized entropy. By mapping secret messages to token probability ranks and dynamically adjusting sampling via context-aware entropy-based adjustments, RTMStega achieves a balance between payload capacity and imperceptibility. Experiments across diverse datasets and models demonstrate that RTMStega triples the payload capacity of mainstream generative steganography, reduces processing time by over 50%, and maintains high text quality, offering a trustworthy solution for secure and efficient covert communication.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A high-capacity linguistic steganography based on entropy-driven rank-token mapping
Jiang, Jun
Zhang, Weiming
Yu, Nenghai
Chen, Kejiang
Cryptography and Security
Artificial Intelligence
Linguistic steganography enables covert communication through embedding secret messages into innocuous texts; however, current methods face critical limitations in payload capacity and security. Traditional modification-based methods introduce detectable anomalies, while retrieval-based strategies suffer from low embedding capacity. Modern generative steganography leverages language models to generate natural stego text but struggles with limited entropy in token predictions, further constraining capacity. To address these issues, we propose an entropy-driven framework called RTMStega that integrates rank-based adaptive coding and context-aware decompression with normalized entropy. By mapping secret messages to token probability ranks and dynamically adjusting sampling via context-aware entropy-based adjustments, RTMStega achieves a balance between payload capacity and imperceptibility. Experiments across diverse datasets and models demonstrate that RTMStega triples the payload capacity of mainstream generative steganography, reduces processing time by over 50%, and maintains high text quality, offering a trustworthy solution for secure and efficient covert communication.
title A high-capacity linguistic steganography based on entropy-driven rank-token mapping
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.23035