Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hartman, Max, Jayaraman, Vidhata, Choraria, Moulik, Savani, Yash, Varshney, Lav R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025)
by: Bagchi, Avi, et al.
Published: (2025)
Context-Gated Associative Retrieval: From Theory to Transformers
by: Choraria, Moulik, et al.
Published: (2026)
by: Choraria, Moulik, et al.
Published: (2026)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
by: Geng, Jianing, et al.
Published: (2025)
by: Geng, Jianing, et al.
Published: (2025)
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
by: Timmer, Roelien C., et al.
Published: (2024)
by: Timmer, Roelien C., et al.
Published: (2024)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
by: Chen, Depeng, et al.
Published: (2024)
by: Chen, Depeng, et al.
Published: (2024)
Asking Back: Interaction-Layer Antidistillation Watermarks
by: Yang, Guang, et al.
Published: (2026)
by: Yang, Guang, et al.
Published: (2026)
Containment Verification: AI Safety Guarantees Independent of Alignment
by: Moon, Royce, et al.
Published: (2026)
by: Moon, Royce, et al.
Published: (2026)
SwitchCIT: Switching for Continual Instruction Tuning
by: Wu, Xinbo, et al.
Published: (2024)
by: Wu, Xinbo, et al.
Published: (2024)
Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images
by: Daw, Arka, et al.
Published: (2024)
by: Daw, Arka, et al.
Published: (2024)
An Approach To Enhance IoT Security In 6G Networks Through Explainable AI
by: Kaur, Navneet, et al.
Published: (2024)
by: Kaur, Navneet, et al.
Published: (2024)
Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
by: Westphal, Charles, et al.
Published: (2026)
by: Westphal, Charles, et al.
Published: (2026)
Hide and Seek: Fingerprinting Large Language Models with Evolutionary Learning
by: Iourovitski, Dmitri, et al.
Published: (2024)
by: Iourovitski, Dmitri, et al.
Published: (2024)
Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models
by: Sinha, Yash, et al.
Published: (2025)
by: Sinha, Yash, et al.
Published: (2025)
Hide Your Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Carrier Articles
by: Wang, Zhilong, et al.
Published: (2024)
by: Wang, Zhilong, et al.
Published: (2024)
ESLD (External Surrogate Latent Defense): A Latent-Space Architecture for Faster, Stronger Prompt-Injection Defense
by: Narendra, Yash
Published: (2026)
by: Narendra, Yash
Published: (2026)
Resource-Aware Deployment Optimization for Collaborative Intrusion Detection in Layered Networks
by: Gómez, André García, et al.
Published: (2026)
by: Gómez, André García, et al.
Published: (2026)
Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
by: Zhu, Delong, et al.
Published: (2024)
by: Zhu, Delong, et al.
Published: (2024)
Hide&Seek: Remove Image Watermarks with Negligible Cost via Pixel-wise Reconstruction
by: Chen, Huajie, et al.
Published: (2026)
by: Chen, Huajie, et al.
Published: (2026)
Energy-Aware Routing to Large Reasoning Models
by: Ellis-Mohr, Austin R., et al.
Published: (2025)
by: Ellis-Mohr, Austin R., et al.
Published: (2025)
Hiding in Plain Sight: An IoT Traffic Camouflage Framework for Enhanced Privacy
by: Worae, Daniel Adu, et al.
Published: (2025)
by: Worae, Daniel Adu, et al.
Published: (2025)
Payload-Aware Intrusion Detection with CMAE and Large Language Models
by: Kim, Yongcheol, et al.
Published: (2025)
by: Kim, Yongcheol, et al.
Published: (2025)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2026)
by: Choudhary, Sarthak, et al.
Published: (2026)
Routing-Aware Explanations for Mixture of Experts Graph Models in Malware Detection
by: Shokouhinejad, Hossein, et al.
Published: (2026)
by: Shokouhinejad, Hossein, et al.
Published: (2026)
Hacking CTFs with Plain Agents
by: Turtayev, Rustem, et al.
Published: (2024)
by: Turtayev, Rustem, et al.
Published: (2024)
SparseJEPA: Sparse Representation Learning of Joint Embedding Predictive Architectures
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
by: Gong, Chen, et al.
Published: (2022)
by: Gong, Chen, et al.
Published: (2022)
Innamark: A Whitespace Replacement Information-Hiding Method
by: Hellmeier, Malte, et al.
Published: (2025)
by: Hellmeier, Malte, et al.
Published: (2025)
Explainability-Aware Evaluation of Transfer Learning Models for IoT DDoS Detection Under Resource Constraints
by: Elsayed, Nelly
Published: (2026)
by: Elsayed, Nelly
Published: (2026)
Eliciting Least-to-Most Reasoning for Phishing URL Detection
by: Trikilis, Holly, et al.
Published: (2026)
by: Trikilis, Holly, et al.
Published: (2026)
Hiding in Plain Sight: Disguising Data Stealing Attacks in Federated Learning
by: Garov, Kostadin, et al.
Published: (2023)
by: Garov, Kostadin, et al.
Published: (2023)
Detecting Various DeFi Price Manipulations with LLM Reasoning
by: Zhong, Juantao, et al.
Published: (2025)
by: Zhong, Juantao, et al.
Published: (2025)
The Hidden Threat in Plain Text: Attacking RAG Data Loaders
by: Castagnaro, Alberto, et al.
Published: (2025)
by: Castagnaro, Alberto, et al.
Published: (2025)
Hiding in Plain Sight: Reframing Hardware Trojan Benchmarking as a Hide&Seek Modification
by: Sarihi, Amin, et al.
Published: (2024)
by: Sarihi, Amin, et al.
Published: (2024)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Red Teaming Large Reasoning Models
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
by: Nie, Yuzhou, et al.
Published: (2025)
by: Nie, Yuzhou, et al.
Published: (2025)
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
by: Chen, Jianhao, et al.
Published: (2025)
by: Chen, Jianhao, et al.
Published: (2025)
Similar Items
-
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025) -
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025) -
Context-Gated Associative Retrieval: From Theory to Transformers
by: Choraria, Moulik, et al.
Published: (2026) -
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
by: Geng, Jianing, et al.
Published: (2025) -
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
by: Timmer, Roelien C., et al.
Published: (2024)