Introspection Adapters: Training LLMs to Report Their Learned Behaviors
Fuente:
arXiv
Saved in:
| Main Authors: | Shenoy, Keshav, Yang, Li, Sheshadri, Abhay, Mindermann, Sören, Lindsey, Jack, Marks, Sam, Wang, Rowan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergent Introspective Awareness in Large Language Models
by: Lindsey, Jack
Published: (2026)
by: Lindsey, Jack
Published: (2026)
The Alignment Problem from a Deep Learning Perspective
by: Ngo, Richard, et al.
Published: (2022)
by: Ngo, Richard, et al.
Published: (2022)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Can LLMs Introspect? A Reality Check
by: Singh, Shashwat, et al.
Published: (2026)
by: Singh, Shashwat, et al.
Published: (2026)
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
by: Hahami, Ely, et al.
Published: (2025)
by: Hahami, Ely, et al.
Published: (2025)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
by: Yang, Diji, et al.
Published: (2024)
by: Yang, Diji, et al.
Published: (2024)
Introspective Diffusion Language Models
by: Yu, Yifan, et al.
Published: (2026)
by: Yu, Yifan, et al.
Published: (2026)
Agentic Misalignment: How LLMs Could Be Insider Threats
by: Lynch, Aengus, et al.
Published: (2025)
by: Lynch, Aengus, et al.
Published: (2025)
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
by: Slocum, Stewart, et al.
Published: (2025)
by: Slocum, Stewart, et al.
Published: (2025)
Can GNN be Good Adapter for LLMs?
by: Huang, Xuanwen, et al.
Published: (2024)
by: Huang, Xuanwen, et al.
Published: (2024)
Position: Introspective Experience from Conversational Environments as a Path to Better Learning
by: Musat, Claudiu Cristian, et al.
Published: (2026)
by: Musat, Claudiu Cristian, et al.
Published: (2026)
Introspection of Thought Helps AI Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
by: Cui, Brandon, et al.
Published: (2026)
by: Cui, Brandon, et al.
Published: (2026)
Adaptive Budget Allocation for Orthogonal-Subspace Adapter Tuning in LLMs Continual Learning
by: Wan, Zhiyi, et al.
Published: (2025)
by: Wan, Zhiyi, et al.
Published: (2025)
Metacognition is all you need? Using Introspection in Generative Agents to Improve Goal-directed Behavior
by: Toy, Jason, et al.
Published: (2024)
by: Toy, Jason, et al.
Published: (2024)
Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection
by: Naphade, Atharv, et al.
Published: (2026)
by: Naphade, Atharv, et al.
Published: (2026)
Emergent Introspection in AI is Content-Agnostic
by: Lederman, Harvey, et al.
Published: (2026)
by: Lederman, Harvey, et al.
Published: (2026)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
Looking Inward: Language Models Can Learn About Themselves by Introspection
by: Binder, Felix J, et al.
Published: (2024)
by: Binder, Felix J, et al.
Published: (2024)
From Simulation to Enaction: Post-trained language models recognize and react to their own generations
by: G., Asvin, et al.
Published: (2026)
by: G., Asvin, et al.
Published: (2026)
Reflect then Learn: Active Prompting for Information Extraction Guided by Introspective Confusion
by: Zhao, Dong, et al.
Published: (2025)
by: Zhao, Dong, et al.
Published: (2025)
Privileged Self-Access Matters for Introspection in AI
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
by: Zala, Abhay, et al.
Published: (2024)
by: Zala, Abhay, et al.
Published: (2024)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
by: Sheshadri, Abhay, et al.
Published: (2026)
by: Sheshadri, Abhay, et al.
Published: (2026)
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
by: Martorell, Nicolas, et al.
Published: (2026)
by: Martorell, Nicolas, et al.
Published: (2026)
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026)
by: Wu, Junfei, et al.
Published: (2026)
Language Models Fail to Introspect About Their Knowledge of Language
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
by: Khullar, Dipika, et al.
Published: (2026)
by: Khullar, Dipika, et al.
Published: (2026)
Sparse Adapter Fusion for Continual Learning in NLP
by: Zeng, Min, et al.
Published: (2026)
by: Zeng, Min, et al.
Published: (2026)
Model Spec Midtraining: Improving How Alignment Training Generalizes
by: Li, Chloe, et al.
Published: (2026)
by: Li, Chloe, et al.
Published: (2026)
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
by: Yin, Ming, et al.
Published: (2025)
by: Yin, Ming, et al.
Published: (2025)
FusionAdapter for Few-Shot Relation Learning in Multimodal Knowledge Graphs
by: Liu, Ran, et al.
Published: (2025)
by: Liu, Ran, et al.
Published: (2025)
Multi-Adapter Representation Interventions via Energy Calibration
by: Yu, Manjiang, et al.
Published: (2026)
by: Yu, Manjiang, et al.
Published: (2026)
Exploration Through Introspection: A Self-Aware Reward Model
by: Petrowski, Michael, et al.
Published: (2026)
by: Petrowski, Michael, et al.
Published: (2026)
Latent Introspection: Models Can Detect Prior Concept Injections
by: Pearson-Vogel, Theia, et al.
Published: (2026)
by: Pearson-Vogel, Theia, et al.
Published: (2026)
Does It Make Sense to Speak of Introspection in Large Language Models?
by: Comsa, Iulia M., et al.
Published: (2025)
by: Comsa, Iulia M., et al.
Published: (2025)
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures
by: Kilic, Ozgur O., et al.
Published: (2025)
by: Kilic, Ozgur O., et al.
Published: (2025)
Similar Items
-
Emergent Introspective Awareness in Large Language Models
by: Lindsey, Jack
Published: (2026) -
The Alignment Problem from a Deep Learning Perspective
by: Ngo, Richard, et al.
Published: (2022) -
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024) -
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
by: Li, Jiaqi, et al.
Published: (2025) -
Can LLMs Introspect? A Reality Check
by: Singh, Shashwat, et al.
Published: (2026)