Can LLMs Introspect? A Reality Check
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Shashwat, Linzen, Tal, Ravfogel, Shauli |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
State over Tokens: Characterizing the Role of Reasoning Tokens
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
Emergence of Linear Truth Encodings in Language Models
by: Ravfogel, Shauli, et al.
Published: (2025)
by: Ravfogel, Shauli, et al.
Published: (2025)
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024)
by: Ravfogel, Shauli, et al.
Published: (2024)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
by: Hong, Yihuai, et al.
Published: (2024)
by: Hong, Yihuai, et al.
Published: (2024)
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
by: Petty, Jackson, et al.
Published: (2026)
by: Petty, Jackson, et al.
Published: (2026)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
by: Petty, Jackson, et al.
Published: (2025)
by: Petty, Jackson, et al.
Published: (2025)
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
by: Hahami, Ely, et al.
Published: (2025)
by: Hahami, Ely, et al.
Published: (2025)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
by: Shenoy, Keshav, et al.
Published: (2026)
by: Shenoy, Keshav, et al.
Published: (2026)
Language Models Struggle to Use Representations Learned In-Context
by: Lepori, Michael A., et al.
Published: (2026)
by: Lepori, Michael A., et al.
Published: (2026)
Rapid Word Learning Through Meta In-Context Learning
by: Wang, Wentao, et al.
Published: (2025)
by: Wang, Wentao, et al.
Published: (2025)
Can LLMs Automate Fact-Checking Article Writing?
by: Sahnan, Dhruv, et al.
Published: (2025)
by: Sahnan, Dhruv, et al.
Published: (2025)
Latent Introspection: Models Can Detect Prior Concept Injections
by: Pearson-Vogel, Theia, et al.
Published: (2026)
by: Pearson-Vogel, Theia, et al.
Published: (2026)
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Looking Inward: Language Models Can Learn About Themselves by Introspection
by: Binder, Felix J, et al.
Published: (2024)
by: Binder, Felix J, et al.
Published: (2024)
Introspective Diffusion Language Models
by: Yu, Yifan, et al.
Published: (2026)
by: Yu, Yifan, et al.
Published: (2026)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
by: Hu, Michael Y., et al.
Published: (2026)
by: Hu, Michael Y., et al.
Published: (2026)
Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
Representation Surgery: Theory and Practice of Affine Steering
by: Singh, Shashwat, et al.
Published: (2024)
by: Singh, Shashwat, et al.
Published: (2024)
SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio
by: Pandey, Satwik, et al.
Published: (2026)
by: Pandey, Satwik, et al.
Published: (2026)
Introspection of Thought Helps AI Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
by: Qiu, Linlu, et al.
Published: (2025)
by: Qiu, Linlu, et al.
Published: (2025)
Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild
by: Yao, Shanle, et al.
Published: (2026)
by: Yao, Shanle, et al.
Published: (2026)
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
by: Eisape, Tiwalayo, et al.
Published: (2023)
by: Eisape, Tiwalayo, et al.
Published: (2023)
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
by: Hagström, Lovisa, et al.
Published: (2024)
by: Hagström, Lovisa, et al.
Published: (2024)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection
by: Naphade, Atharv, et al.
Published: (2026)
by: Naphade, Atharv, et al.
Published: (2026)
Can LLMs Capture Human Preferences?
by: Goli, Ali, et al.
Published: (2023)
by: Goli, Ali, et al.
Published: (2023)
Emergent Introspection in AI is Content-Agnostic
by: Lederman, Harvey, et al.
Published: (2026)
by: Lederman, Harvey, et al.
Published: (2026)
Deep Active Learning: A Reality Check
by: Gashi, Edrina, et al.
Published: (2024)
by: Gashi, Edrina, et al.
Published: (2024)
InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs
by: Bai, Yuzhuo, et al.
Published: (2026)
by: Bai, Yuzhuo, et al.
Published: (2026)
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
by: Sinha, Akshit, et al.
Published: (2025)
by: Sinha, Akshit, et al.
Published: (2025)
Emergent Introspective Awareness in Large Language Models
by: Lindsey, Jack
Published: (2026)
by: Lindsey, Jack
Published: (2026)
Privileged Self-Access Matters for Introspection in AI
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Exploration Through Introspection: A Self-Aware Reward Model
by: Petrowski, Michael, et al.
Published: (2026)
by: Petrowski, Michael, et al.
Published: (2026)
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
by: Martorell, Nicolas, et al.
Published: (2026)
by: Martorell, Nicolas, et al.
Published: (2026)
Language Models Fail to Introspect About Their Knowledge of Language
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Preserving Task-Relevant Information Under Linear Concept Removal
by: Holstege, Floris, et al.
Published: (2025)
by: Holstege, Floris, et al.
Published: (2025)
Log-linear Guardedness and its Implications
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
How LLMs Fail to Support Fact-Checking
by: Proma, Adiba Mahbub, et al.
Published: (2025)
by: Proma, Adiba Mahbub, et al.
Published: (2025)
Behavioral Cloning Models Reality Check for Autonomous Driving
by: Yildirim, Mustafa, et al.
Published: (2024)
by: Yildirim, Mustafa, et al.
Published: (2024)
Similar Items
-
State over Tokens: Characterizing the Role of Reasoning Tokens
by: Levy, Mosh, et al.
Published: (2025) -
Emergence of Linear Truth Encodings in Language Models
by: Ravfogel, Shauli, et al.
Published: (2025) -
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024) -
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
by: Hong, Yihuai, et al.
Published: (2024) -
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
by: Petty, Jackson, et al.
Published: (2026)