Hodoscope: Unsupervised Monitoring for AI Misbehaviors
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Ziqian, Saxena, Shashwat, Raghunathan, Aditi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026)
by: Bercovich, Ivan, et al.
Published: (2026)
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
by: Zhong, Ziqian, et al.
Published: (2026)
by: Zhong, Ziqian, et al.
Published: (2026)
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
by: Baker, Bowen, et al.
Published: (2025)
by: Baker, Bowen, et al.
Published: (2025)
Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
by: Turtayev, Rustem, et al.
Published: (2025)
by: Turtayev, Rustem, et al.
Published: (2025)
Mitigating Modal Imbalance in Multimodal Reasoning
by: Wu, Chen Henry, et al.
Published: (2025)
by: Wu, Chen Henry, et al.
Published: (2025)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
Training Agents to Self-Report Misbehavior
by: Lee, Bruce W., et al.
Published: (2026)
by: Lee, Bruce W., et al.
Published: (2026)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
Reasoning as an Adaptive Defense for Safety
by: Kim, Taeyoun, et al.
Published: (2025)
by: Kim, Taeyoun, et al.
Published: (2025)
LLMScan: Causal Scan for LLM Misbehavior Detection
by: Zhang, Mengdi, et al.
Published: (2024)
by: Zhang, Mengdi, et al.
Published: (2024)
Mode-Conditioning Unlocks Superior Test-Time Scaling
by: Wu, Chen Henry, et al.
Published: (2025)
by: Wu, Chen Henry, et al.
Published: (2025)
GradientSpace: Unsupervised Data Clustering for Improved Instruction Tuning
by: Sridharan, Shrihari, et al.
Published: (2025)
by: Sridharan, Shrihari, et al.
Published: (2025)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Wink: Recovering from Misbehaviors in Coding Agents
by: Nanda, Rahul, et al.
Published: (2026)
by: Nanda, Rahul, et al.
Published: (2026)
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025)
by: Dang, Xingyu, et al.
Published: (2025)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
by: Min, Nay Myat, et al.
Published: (2026)
by: Min, Nay Myat, et al.
Published: (2026)
A Novel Labeled Human Voice Signal Dataset for Misbehavior Detection
by: Raza, Ali, et al.
Published: (2024)
by: Raza, Ali, et al.
Published: (2024)
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
by: Kim, Eungyeup, et al.
Published: (2023)
by: Kim, Eungyeup, et al.
Published: (2023)
Synthetic Data for Robust AI Model Development in Regulated Enterprises
by: Godbole, Aditi
Published: (2025)
by: Godbole, Aditi
Published: (2025)
AttentionGuard: Transformer-based Misbehavior Detection for Secure Vehicular Platoons
by: Li, Hexu, et al.
Published: (2025)
by: Li, Hexu, et al.
Published: (2025)
Algorithmic Capabilities of Random Transformers
by: Zhong, Ziqian, et al.
Published: (2024)
by: Zhong, Ziqian, et al.
Published: (2024)
S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
by: Chavan, Arnav, et al.
Published: (2026)
by: Chavan, Arnav, et al.
Published: (2026)
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
by: Gai, Jingchu, et al.
Published: (2026)
by: Gai, Jingchu, et al.
Published: (2026)
Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection
by: Kalogiannis, Konstantinos, et al.
Published: (2025)
by: Kalogiannis, Konstantinos, et al.
Published: (2025)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
by: Vishwanathan, Manoj, et al.
Published: (2026)
by: Vishwanathan, Manoj, et al.
Published: (2026)
SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio
by: Pandey, Satwik, et al.
Published: (2026)
by: Pandey, Satwik, et al.
Published: (2026)
Proper Scoring Rules for Agentic Uncertainty Quantification
by: Raghu, Suresh, et al.
Published: (2026)
by: Raghu, Suresh, et al.
Published: (2026)
Can LLMs Introspect? A Reality Check
by: Singh, Shashwat, et al.
Published: (2026)
by: Singh, Shashwat, et al.
Published: (2026)
A3D: Agentic AI flow for autonomous Accelerator Design
by: Nallathambi, Abinand, et al.
Published: (2026)
by: Nallathambi, Abinand, et al.
Published: (2026)
From Horizontal Layering to Vertical Integration: A Comparative Study of the AI-Driven Software Development Paradigm
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
AI based Content Creation and Product Recommendation Applications in E-commerce: An Ethical overview
by: Jain, Aditi Madhusudan, et al.
Published: (2025)
by: Jain, Aditi Madhusudan, et al.
Published: (2025)
A Fourth Wave of Open Data? Exploring the Spectrum of Scenarios for Open Data and Generative AI
by: Chafetz, Hannah, et al.
Published: (2024)
by: Chafetz, Hannah, et al.
Published: (2024)
PAMPOS: Causal Transformer-based Trajectory Prediction for Attack-Agnostic Misbehavior Detection in V2X Networks
by: Kalogiannis, Konstantinos, et al.
Published: (2026)
by: Kalogiannis, Konstantinos, et al.
Published: (2026)
Towards an AI Observatory for the Nuclear Sector: A tool for anticipatory governance
by: Verma, Aditi, et al.
Published: (2025)
by: Verma, Aditi, et al.
Published: (2025)
Overtrained Language Models Are Harder to Fine-Tune
by: Springer, Jacob Mitchell, et al.
Published: (2025)
by: Springer, Jacob Mitchell, et al.
Published: (2025)
Generative AI in Ship Design
by: Thakur, Sahil, et al.
Published: (2024)
by: Thakur, Sahil, et al.
Published: (2024)
Similar Items
-
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026) -
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025) -
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
by: Zhong, Ziqian, et al.
Published: (2025) -
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026) -
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
by: Zhong, Ziqian, et al.
Published: (2026)