Why Do Language Model Agents Whistleblow?
Fuente:
arXiv
Saved in:
| Main Authors: | Agrawal, Kushal, Xiao, Frank, Bergman, Guido, Stickland, Asa Cooper |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
Neural Diversity Regularizes Hallucinations in Language Models
by: Chakrabarti, Kushal, et al.
Published: (2025)
by: Chakrabarti, Kushal, et al.
Published: (2025)
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
by: Agrawal, Saurabh, et al.
Published: (2025)
by: Agrawal, Saurabh, et al.
Published: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Why Larger Language Models Do In-context Learning Differently?
by: Shi, Zhenmei, et al.
Published: (2024)
by: Shi, Zhenmei, et al.
Published: (2024)
Don't Waste Your Time: Early Stopping Cross-Validation
by: Bergman, Edward, et al.
Published: (2024)
by: Bergman, Edward, et al.
Published: (2024)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
by: Kang, Deokhyung, et al.
Published: (2025)
by: Kang, Deokhyung, et al.
Published: (2025)
Why Do Unlearnable Examples Work: A Novel Perspective of Mutual Information
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Inference Energy and Latency in AI-Mediated Education: A Learning-per-Watt Analysis of Edge and Cloud Models
by: Khemani, Kushal
Published: (2026)
by: Khemani, Kushal
Published: (2026)
DoWhy-GCM: An extension of DoWhy for causal inference in graphical causal models
by: Blöbaum, Patrick, et al.
Published: (2022)
by: Blöbaum, Patrick, et al.
Published: (2022)
Why Do Time Series Models Need Long Context Windows?
by: Butera, Luca, et al.
Published: (2026)
by: Butera, Luca, et al.
Published: (2026)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Evaluating LLM Agent Collusion in Double Auctions
by: Agrawal, Kushal, et al.
Published: (2025)
by: Agrawal, Kushal, et al.
Published: (2025)
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
by: Stickland, Asa Cooper, et al.
Published: (2024)
by: Stickland, Asa Cooper, et al.
Published: (2024)
Fast Benchmarking of Asynchronous Multi-Fidelity Optimization on Zero-Cost Benchmarks
by: Watanabe, Shuhei, et al.
Published: (2024)
by: Watanabe, Shuhei, et al.
Published: (2024)
Analyzing Memorization in Large Language Models through the Lens of Model Attribution
by: Menta, Tarun Ram, et al.
Published: (2025)
by: Menta, Tarun Ram, et al.
Published: (2025)
Unraveling the cognitive patterns of Large Language Models through module communities
by: Bhandari, Kushal Raj, et al.
Published: (2025)
by: Bhandari, Kushal Raj, et al.
Published: (2025)
Why Do Transformers Fail to Forecast Time Series In-Context?
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
Why Do Some Inputs Break Low-Bit LLM Quantization?
by: Chang, Ting-Yun, et al.
Published: (2025)
by: Chang, Ting-Yun, et al.
Published: (2025)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
Why Do Neural Networks Forget: A Study of Collapse in Continual Learning
by: Zhu, Yunqin, et al.
Published: (2026)
by: Zhu, Yunqin, et al.
Published: (2026)
Zero-shot Meta-learning for Tabular Prediction Tasks with Adversarially Pre-trained Transformer
by: Wu, Yulun, et al.
Published: (2025)
by: Wu, Yulun, et al.
Published: (2025)
Learnware of Language Models: Specialized Small Language Models Can Do Big
by: Tan, Zhi-Hao, et al.
Published: (2025)
by: Tan, Zhi-Hao, et al.
Published: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
Do Large Language Models (LLMs) Understand Chronology?
by: Wongchamcharoen, Pattaraphon Kenny, et al.
Published: (2025)
by: Wongchamcharoen, Pattaraphon Kenny, et al.
Published: (2025)
DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents
by: Zhao, Jiahao, et al.
Published: (2026)
by: Zhao, Jiahao, et al.
Published: (2026)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
by: Sharma, Archit, et al.
Published: (2024)
by: Sharma, Archit, et al.
Published: (2024)
Why the Agent Made that Decision: Contrastive Explanation Learning for Reinforcement Learning
by: Zuo, Rui, et al.
Published: (2024)
by: Zuo, Rui, et al.
Published: (2024)
Why Representation Engineering Works: A Theoretical and Empirical Study in Vision-Language Models
by: Tian, Bowei, et al.
Published: (2025)
by: Tian, Bowei, et al.
Published: (2025)
World Modelling Improves Language Model Agents
by: Guo, Shangmin, et al.
Published: (2025)
by: Guo, Shangmin, et al.
Published: (2025)
Math Takes Two: A test for emergent mathematical reasoning in communication
by: Cooper, Michael, et al.
Published: (2026)
by: Cooper, Michael, et al.
Published: (2026)
Estimating the Empowerment of Language Model Agents
by: Song, Jinyeop, et al.
Published: (2025)
by: Song, Jinyeop, et al.
Published: (2025)
A Methodology Establishing Linear Convergence of Adaptive Gradient Methods under PL Inequality
by: Chakrabarti, Kushal, et al.
Published: (2024)
by: Chakrabarti, Kushal, et al.
Published: (2024)
Adaptive Few-Shot Learning (AFSL): Tackling Data Scarcity with Stability, Robustness, and Versatility
by: Agrawal, Rishabh
Published: (2025)
by: Agrawal, Rishabh
Published: (2025)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
HiCL: Hippocampal-Inspired Continual Learning
by: Kapoor, Kushal, et al.
Published: (2025)
by: Kapoor, Kushal, et al.
Published: (2025)
Controllable Game Level Generation: Assessing the Effect of Negative Examples in GAN Models
by: Bazzaz, Mahsa, et al.
Published: (2024)
by: Bazzaz, Mahsa, et al.
Published: (2024)
The Five Ws of Multi-Agent Communication: Who Talks to Whom, When, What, and Why -- A Survey from MARL to Emergent Language and LLMs
by: Chen, Jingdi, et al.
Published: (2026)
by: Chen, Jingdi, et al.
Published: (2026)
Aligning Agents like Large Language Models
by: Jelley, Adam, et al.
Published: (2024)
by: Jelley, Adam, et al.
Published: (2024)
Similar Items
-
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023) -
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024) -
Neural Diversity Regularizes Hallucinations in Language Models
by: Chakrabarti, Kushal, et al.
Published: (2025) -
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
by: Agrawal, Saurabh, et al.
Published: (2025) -
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)