Among Us: A Sandbox for Measuring and Detecting Agentic Deception
Fuente:
arXiv
Saved in:
| Main Authors: | Golechha, Satvik, Garriga-Alonso, Adrià |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progress Measures for Grokking on Real-world Tasks
by: Golechha, Satvik
Published: (2024)
by: Golechha, Satvik
Published: (2024)
Building Better Deception Probes Using Targeted Instruction Pairs
by: Natarajan, Vikram, et al.
Published: (2026)
by: Natarajan, Vikram, et al.
Published: (2026)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Training Neural Networks for Modularity aids Interpretability
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
by: Chanin, David, et al.
Published: (2026)
by: Chanin, David, et al.
Published: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
NICE: To Optimize In-Context Examples or Not?
by: Srivastava, Pragya, et al.
Published: (2024)
by: Srivastava, Pragya, et al.
Published: (2024)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
by: Arcuschin, Iván, et al.
Published: (2026)
by: Arcuschin, Iván, et al.
Published: (2026)
Studying Cross-cluster Modularity in Neural Networks
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
by: Lidayan, Aly, et al.
Published: (2025)
by: Lidayan, Aly, et al.
Published: (2025)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
by: Taufeeque, Mohammad, et al.
Published: (2025)
by: Taufeeque, Mohammad, et al.
Published: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
by: Bush, Thomas, et al.
Published: (2025)
by: Bush, Thomas, et al.
Published: (2025)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
by: Balappanawar, Ishwar, et al.
Published: (2025)
by: Balappanawar, Ishwar, et al.
Published: (2025)
DiFR: Inference Verification Despite Nondeterminism
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Planning in a recurrent neural network that plays Sokoban
by: Taufeeque, Mohammad, et al.
Published: (2024)
by: Taufeeque, Mohammad, et al.
Published: (2024)
Hypothesis Testing the Circuit Hypothesis in LLMs
by: Shi, Claudia, et al.
Published: (2024)
by: Shi, Claudia, et al.
Published: (2024)
CantorNet: A Sandbox for Testing Geometrical and Topological Complexity Measures
by: Lewandowski, Michal, et al.
Published: (2024)
by: Lewandowski, Michal, et al.
Published: (2024)
Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method
by: Tripathi, Satvik
Published: (2025)
by: Tripathi, Satvik
Published: (2025)
Investigating the Indirect Object Identification circuit in Mamba
by: Ensign, Danielle, et al.
Published: (2024)
by: Ensign, Danielle, et al.
Published: (2024)
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
by: Ravindran, Santhosh Kumar
Published: (2025)
by: Ravindran, Santhosh Kumar
Published: (2025)
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
by: Vishwakarma, Harsh, et al.
Published: (2025)
by: Vishwakarma, Harsh, et al.
Published: (2025)
Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi
by: Rustagi, Arnav, et al.
Published: (2025)
by: Rustagi, Arnav, et al.
Published: (2025)
An Introductory Survey to Autoencoder-based Deep Clustering -- Sandboxes for Combining Clustering with Deep Learning
by: Leiber, Collin, et al.
Published: (2025)
by: Leiber, Collin, et al.
Published: (2025)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
Deceptive Exploration in Multi-armed Bandits
by: Vurankaya, I. Arda, et al.
Published: (2025)
by: Vurankaya, I. Arda, et al.
Published: (2025)
Enhancing Malware Detection by Integrating Machine Learning with Cuckoo Sandbox
by: Alshmarni, Amaal F., et al.
Published: (2023)
by: Alshmarni, Amaal F., et al.
Published: (2023)
Information-theoretic Distinctions Between Deception and Confusion
by: Young, Robin
Published: (2025)
by: Young, Robin
Published: (2025)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
by: Zhao, Zhenyu, et al.
Published: (2026)
by: Zhao, Zhenyu, et al.
Published: (2026)
Decoding Answers Before Chain-of-Thought: Evidence from Pre-CoT Probes and Activation Steering
by: Cox, Kyle, et al.
Published: (2026)
by: Cox, Kyle, et al.
Published: (2026)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
When Truthful Representations Flip Under Deceptive Instructions?
by: Long, Xianxuan, et al.
Published: (2025)
by: Long, Xianxuan, et al.
Published: (2025)
MemER: Scaling Up Memory for Robot Control via Experience Retrieval
by: Sridhar, Ajay, et al.
Published: (2025)
by: Sridhar, Ajay, et al.
Published: (2025)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
by: Lee, Hayeong, et al.
Published: (2026)
by: Lee, Hayeong, et al.
Published: (2026)
DecepChain: Inducing Deceptive Reasoning in Large Language Models
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
by: Williams, Marcus, et al.
Published: (2024)
by: Williams, Marcus, et al.
Published: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Data Leakage and Deceptive Performance: A Critical Examination of Credit Card Fraud Detection Methodologies
by: Hayat, Khizar, et al.
Published: (2025)
by: Hayat, Khizar, et al.
Published: (2025)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
Similar Items
-
Progress Measures for Grokking on Real-world Tasks
by: Golechha, Satvik
Published: (2024) -
Building Better Deception Probes Using Targeted Instruction Pairs
by: Natarajan, Vikram, et al.
Published: (2026) -
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024) -
Training Neural Networks for Modularity aids Interpretability
by: Golechha, Satvik, et al.
Published: (2024) -
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
by: Chanin, David, et al.
Published: (2026)