Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chloe, Phuong, Mary, Tan, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Capability Self-Assessment: Teaching LLMs to Know Their Limits
by: Yang, Haoyan, et al.
Published: (2026)
by: Yang, Haoyan, et al.
Published: (2026)
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
by: Jotautaite, Monika, et al.
Published: (2025)
by: Jotautaite, Monika, et al.
Published: (2025)
Spilled Energy in Large Language Models
by: Minut, Adrian Robert, et al.
Published: (2026)
by: Minut, Adrian Robert, et al.
Published: (2026)
Tricking LLM-Based NPCs into Spilling Secrets
by: Shiomi, Kyohei, et al.
Published: (2025)
by: Shiomi, Kyohei, et al.
Published: (2025)
Teaching LLMs to Ask: Self-Querying Category-Theoretic Planning for Under-Specified Reasoning
by: Qu, Shuhui
Published: (2026)
by: Qu, Shuhui
Published: (2026)
On the Hidden Objective Biases of Group-based Reinforcement Learning
by: Fontana, Aleksandar, et al.
Published: (2026)
by: Fontana, Aleksandar, et al.
Published: (2026)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
by: Xu, Tianyang, et al.
Published: (2024)
by: Xu, Tianyang, et al.
Published: (2024)
Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs
by: Yan, Qianqi, et al.
Published: (2025)
by: Yan, Qianqi, et al.
Published: (2025)
MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures
by: Zhang, Shan, et al.
Published: (2026)
by: Zhang, Shan, et al.
Published: (2026)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
by: Li, Mingjie, et al.
Published: (2026)
by: Li, Mingjie, et al.
Published: (2026)
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
by: Yu, Jiahao, et al.
Published: (2024)
by: Yu, Jiahao, et al.
Published: (2024)
When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
by: Wei, Jiankun, et al.
Published: (2024)
by: Wei, Jiankun, et al.
Published: (2024)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)
by: Lu, Yu-An, et al.
Published: (2026)
Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
by: Li, Xinran, et al.
Published: (2025)
by: Li, Xinran, et al.
Published: (2025)
Learning to Plan Before Answering: Self-Teaching LLMs to Learn Abstract Plans for Problem Solving
by: Zhang, Jin, et al.
Published: (2025)
by: Zhang, Jin, et al.
Published: (2025)
The Challenge of Teaching Reasoning to LLMs Without RL or Distillation
by: Du, Wei, et al.
Published: (2025)
by: Du, Wei, et al.
Published: (2025)
Oil Spill Segmentation using Deep Encoder-Decoder models
by: Satyanarayana, Abhishek Ramanathapura, et al.
Published: (2023)
by: Satyanarayana, Abhishek Ramanathapura, et al.
Published: (2023)
Self-Routing: Parameter-Free Expert Routing from Hidden States
by: Mohamud, Jama Hussein, et al.
Published: (2026)
by: Mohamud, Jama Hussein, et al.
Published: (2026)
Measuring Teaching with LLMs
by: Hardy, Michael
Published: (2025)
by: Hardy, Michael
Published: (2025)
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
by: Bianchi, Owen, et al.
Published: (2025)
by: Bianchi, Owen, et al.
Published: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
by: Yuan, Huining, et al.
Published: (2025)
by: Yuan, Huining, et al.
Published: (2025)
Estimating the Self-Consistency of LLMs
by: Nowak, Robert
Published: (2025)
by: Nowak, Robert
Published: (2025)
Structure Enables Effective Self-Localization of Errors in LLMs
by: Samanta, Ankur, et al.
Published: (2026)
by: Samanta, Ankur, et al.
Published: (2026)
Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning
by: Draganov, Andrew, et al.
Published: (2026)
by: Draganov, Andrew, et al.
Published: (2026)
CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
MEMENTO: Teaching LLMs to Manage Their Own Context
by: Kontonis, Vasilis, et al.
Published: (2026)
by: Kontonis, Vasilis, et al.
Published: (2026)
From Solving to Verifying: A Unified Objective for Robust Reasoning in LLMs
by: Wang, Xiaoxuan, et al.
Published: (2025)
by: Wang, Xiaoxuan, et al.
Published: (2025)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
by: Polowczyk, Agnieszka, et al.
Published: (2025)
by: Polowczyk, Agnieszka, et al.
Published: (2025)
Collaborative Expert LLMs Guided Multi-Objective Molecular Optimization
by: Yu, Jiajun, et al.
Published: (2025)
by: Yu, Jiajun, et al.
Published: (2025)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
by: Shenoy, Keshav, et al.
Published: (2026)
by: Shenoy, Keshav, et al.
Published: (2026)
SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks
by: Hu, Jinchao, et al.
Published: (2026)
by: Hu, Jinchao, et al.
Published: (2026)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
Similar Items
-
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025) -
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026) -
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
by: Adiletta, Andrew, et al.
Published: (2025) -
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024) -
Capability Self-Assessment: Teaching LLMs to Know Their Limits
by: Yang, Haoyan, et al.
Published: (2026)