Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Chrabąszcz, Maciej, Szymczyk, Aleksander, Sendera, Marcin, Trzciński, Tomasz, Cygert, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient LLM Moderation with Multi-Layer Latent Prototypes
by: Chrabąszcz, Maciej, et al.
Published: (2025)
by: Chrabąszcz, Maciej, et al.
Published: (2025)
Conditioned Activation Transport for T2I Safety Steering
by: Chrabąszcz, Maciej, et al.
Published: (2026)
by: Chrabąszcz, Maciej, et al.
Published: (2026)
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026)
by: Meng, Wenlong, et al.
Published: (2026)
Conti Inc.: Understanding the Internal Discussions of a large Ransomware-as-a-Service Operator with Machine Learning
by: Ruellan, Estelle, et al.
Published: (2023)
by: Ruellan, Estelle, et al.
Published: (2023)
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
by: Chuang, Ko-Wei, et al.
Published: (2025)
by: Chuang, Ko-Wei, et al.
Published: (2025)
Expert Selections In MoE Models Reveal (Almost) As Much As Text
by: Nuriyev, Amir, et al.
Published: (2026)
by: Nuriyev, Amir, et al.
Published: (2026)
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
by: He, Xuanli, et al.
Published: (2026)
by: He, Xuanli, et al.
Published: (2026)
garak: A Framework for Security Probing Large Language Models
by: Derczynski, Leon, et al.
Published: (2024)
by: Derczynski, Leon, et al.
Published: (2024)
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2026)
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2026)
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
by: Sandoval, Aaron, et al.
Published: (2025)
by: Sandoval, Aaron, et al.
Published: (2025)
Subversion via Focal Points: Investigating Collusion in LLM Monitoring
by: Järviniemi, Olli
Published: (2025)
by: Järviniemi, Olli
Published: (2025)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
by: Zhou, Yukai, et al.
Published: (2025)
by: Zhou, Yukai, et al.
Published: (2025)
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs
by: Shaaban, Mohamed, et al.
Published: (2026)
by: Shaaban, Mohamed, et al.
Published: (2026)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
by: Zhao, Gejian, et al.
Published: (2025)
by: Zhao, Gejian, et al.
Published: (2025)
BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
by: Luo, Xuan, et al.
Published: (2026)
by: Luo, Xuan, et al.
Published: (2026)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
Internal Safety Collapse in Frontier Large Language Models
by: Wu, Yutao, et al.
Published: (2026)
by: Wu, Yutao, et al.
Published: (2026)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
From Retrieval to Reasoning: A Framework for Cyber Threat Intelligence NER with Explicit and Adaptive Instructions
by: Peng, Jiaren, et al.
Published: (2025)
by: Peng, Jiaren, et al.
Published: (2025)
Adversarial Robustness through Dynamic Ensemble Learning
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
Adversarial Text Generation with Dynamic Contextual Perturbation
by: Waghela, Hetvi, et al.
Published: (2025)
by: Waghela, Hetvi, et al.
Published: (2025)
Provable Secure Steganography Based on Adaptive Dynamic Sampling
by: Pang, Kaiyi, et al.
Published: (2025)
by: Pang, Kaiyi, et al.
Published: (2025)
LRCTI: A Large Language Model-Based Framework for Multi-Step Evidence Retrieval and Reasoning in Cyber Threat Intelligence Credibility Verification
by: Tang, Fengxiao, et al.
Published: (2025)
by: Tang, Fengxiao, et al.
Published: (2025)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank Compression
by: Yao, Kai, et al.
Published: (2024)
by: Yao, Kai, et al.
Published: (2024)
TempCharBERT: Keystroke Dynamics for Continuous Access Control Based on Pre-trained Language Models
by: Simão, Matheus, et al.
Published: (2024)
by: Simão, Matheus, et al.
Published: (2024)
Sentinels of the Stream: Unleashing Large Language Models for Dynamic Packet Classification in Software Defined Networks -- Position Paper
by: Murtuza, Shariq
Published: (2024)
by: Murtuza, Shariq
Published: (2024)
$PD^3F$: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
by: Ahmed, Nesreen K., et al.
Published: (2026)
by: Ahmed, Nesreen K., et al.
Published: (2026)
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
by: Berezin, Sergey, et al.
Published: (2025)
by: Berezin, Sergey, et al.
Published: (2025)
Measuring Copyright Risks of Large Language Model via Partial Information Probing
by: Zhao, Weijie, et al.
Published: (2024)
by: Zhao, Weijie, et al.
Published: (2024)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Large Reasoning Models Are Autonomous Jailbreak Agents
by: Hagendorff, Thilo, et al.
Published: (2025)
by: Hagendorff, Thilo, et al.
Published: (2025)
Similar Items
-
Efficient LLM Moderation with Multi-Layer Latent Prototypes
by: Chrabąszcz, Maciej, et al.
Published: (2025) -
Conditioned Activation Transport for T2I Safety Steering
by: Chrabąszcz, Maciej, et al.
Published: (2026) -
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026) -
Conti Inc.: Understanding the Internal Discussions of a large Ransomware-as-a-Service Operator with Machine Learning
by: Ruellan, Estelle, et al.
Published: (2023) -
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
by: Chuang, Ko-Wei, et al.
Published: (2025)