Building Better Deception Probes Using Targeted Instruction Pairs
Fuente:
arXiv
Saved in:
| Main Authors: | Natarajan, Vikram, Jain, Devina, Arora, Shivam, Golechha, Satvik, Bloom, Joseph |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Progress Measures for Grokking on Real-world Tasks
by: Golechha, Satvik
Published: (2024)
by: Golechha, Satvik
Published: (2024)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Training Neural Networks for Modularity aids Interpretability
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
NICE: To Optimize In-Context Examples or Not?
by: Srivastava, Pragya, et al.
Published: (2024)
by: Srivastava, Pragya, et al.
Published: (2024)
Studying Cross-cluster Modularity in Neural Networks
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
by: Lidayan, Aly, et al.
Published: (2025)
by: Lidayan, Aly, et al.
Published: (2025)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
by: Balappanawar, Ishwar, et al.
Published: (2025)
by: Balappanawar, Ishwar, et al.
Published: (2025)
Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment
by: Kitkana, Chayanon, et al.
Published: (2026)
by: Kitkana, Chayanon, et al.
Published: (2026)
ContextBench: Modifying Contexts for Targeted Latent Activation
by: Graham, Robert, et al.
Published: (2025)
by: Graham, Robert, et al.
Published: (2025)
When Truthful Representations Flip Under Deceptive Instructions?
by: Long, Xianxuan, et al.
Published: (2025)
by: Long, Xianxuan, et al.
Published: (2025)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2024)
by: Chanin, David, et al.
Published: (2024)
Building Expressive and Tractable Probabilistic Generative Models: A Review
by: Sidheekh, Sahil, et al.
Published: (2024)
by: Sidheekh, Sahil, et al.
Published: (2024)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
by: Taufeeque, Mohammad, et al.
Published: (2026)
by: Taufeeque, Mohammad, et al.
Published: (2026)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
by: Williams, Marcus, et al.
Published: (2024)
by: Williams, Marcus, et al.
Published: (2024)
Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection
by: Ryser, Alain, et al.
Published: (2024)
by: Ryser, Alain, et al.
Published: (2024)
Skill-Targeted Adaptive Training
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method
by: Tripathi, Satvik
Published: (2025)
by: Tripathi, Satvik
Published: (2025)
Better Alignment with Instruction Back-and-Forth Translation
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
Deceptive Exploration in Multi-armed Bandits
by: Vurankaya, I. Arda, et al.
Published: (2025)
by: Vurankaya, I. Arda, et al.
Published: (2025)
Instruction-tuned Language Models are Better Knowledge Learners
by: Jiang, Zhengbao, et al.
Published: (2024)
by: Jiang, Zhengbao, et al.
Published: (2024)
Optimal Policy Sparsification and Low Rank Decomposition for Deep Reinforcement Learning
by: Goddla, Vikram
Published: (2024)
by: Goddla, Vikram
Published: (2024)
Information-theoretic Distinctions Between Deception and Confusion
by: Young, Robin
Published: (2025)
by: Young, Robin
Published: (2025)
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026)
by: Kramár, János, et al.
Published: (2026)
ProgRM: Build Better GUI Agents with Progress Rewards
by: Zhang, Danyang, et al.
Published: (2025)
by: Zhang, Danyang, et al.
Published: (2025)
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
by: Dawes, Cutter, et al.
Published: (2026)
by: Dawes, Cutter, et al.
Published: (2026)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
MemER: Scaling Up Memory for Robot Control via Experience Retrieval
by: Sridhar, Ajay, et al.
Published: (2025)
by: Sridhar, Ajay, et al.
Published: (2025)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
by: Kumar, Sachin
Published: (2026)
by: Kumar, Sachin
Published: (2026)
Interacting Large Language Model Agents. Interpretable Models and Social Learning
by: Jain, Adit, et al.
Published: (2024)
by: Jain, Adit, et al.
Published: (2024)
Geometry-Aware Probabilistic Circuits via Voronoi Tessellations
by: Sidheekh, Sahil, et al.
Published: (2026)
by: Sidheekh, Sahil, et al.
Published: (2026)
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024)
by: Wan, Xingchen, et al.
Published: (2024)
DecepChain: Inducing Deceptive Reasoning in Large Language Models
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
Mesh-based Super-resolution of Detonation Flows with Multiscale Graph Transformers
by: Barwey, Shivam, et al.
Published: (2025)
by: Barwey, Shivam, et al.
Published: (2025)
Beyond Winning: Margin of Victory Relative to Expectation Unlocks Accurate Skill Ratings
by: Shorewala, Shivam, et al.
Published: (2025)
by: Shorewala, Shivam, et al.
Published: (2025)
Intrinsic Dimension Estimation for Radio Galaxy Zoo using Diffusion Models
by: Roset, Joan Font-Quer, et al.
Published: (2025)
by: Roset, Joan Font-Quer, et al.
Published: (2025)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
by: Wanaskar, Kapil, et al.
Published: (2026)
by: Wanaskar, Kapil, et al.
Published: (2026)
Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
Similar Items
-
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025) -
Progress Measures for Grokking on Real-world Tasks
by: Golechha, Satvik
Published: (2024) -
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024) -
Training Neural Networks for Modularity aids Interpretability
by: Golechha, Satvik, et al.
Published: (2024) -
NICE: To Optimize In-Context Examples or Not?
by: Srivastava, Pragya, et al.
Published: (2024)