Emergent Introspection in AI is Content-Agnostic
Fuente:
arXiv
Saved in:
| Main Authors: | Lederman, Harvey, Mahowald, Kyle |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Privileged Self-Access Matters for Introspection in AI
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Language Models Fail to Introspect About Their Knowledge of Language
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Are Language Models More Like Libraries or Like Librarians? Bibliotechnism, the Novel Reference Problem, and the Attitudes of LLMs
by: Lederman, Harvey, et al.
Published: (2024)
by: Lederman, Harvey, et al.
Published: (2024)
The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models
by: Drucker, Daniel, et al.
Published: (2026)
by: Drucker, Daniel, et al.
Published: (2026)
Emergent Introspective Awareness in Large Language Models
by: Lindsey, Jack
Published: (2026)
by: Lindsey, Jack
Published: (2026)
Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
by: Misra, Kanishka, et al.
Published: (2024)
by: Misra, Kanishka, et al.
Published: (2024)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
by: Rozner, Josh, et al.
Published: (2021)
by: Rozner, Josh, et al.
Published: (2021)
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
by: Boguraev, Sasha, et al.
Published: (2025)
by: Boguraev, Sasha, et al.
Published: (2025)
semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces
by: Ranganathan, Jwalanthi, et al.
Published: (2025)
by: Ranganathan, Jwalanthi, et al.
Published: (2025)
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
by: Boguraev, Sasha, et al.
Published: (2024)
by: Boguraev, Sasha, et al.
Published: (2024)
Language models align with human judgments on key grammatical constructions
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
by: Nemitz, Jonathan, et al.
Published: (2026)
by: Nemitz, Jonathan, et al.
Published: (2026)
What Can String Probability Tell Us About Grammaticality?
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
by: Sheffield, William, et al.
Published: (2025)
by: Sheffield, William, et al.
Published: (2025)
Mission: Impossible Language Models
by: Kallini, Julie, et al.
Published: (2024)
by: Kallini, Julie, et al.
Published: (2024)
Does It Make Sense to Speak of Introspection in Large Language Models?
by: Comsa, Iulia M., et al.
Published: (2025)
by: Comsa, Iulia M., et al.
Published: (2025)
Dissociating language and thought in large language models
by: Mahowald, Kyle, et al.
Published: (2023)
by: Mahowald, Kyle, et al.
Published: (2023)
Looking Inward: Language Models Can Learn About Themselves by Introspection
by: Binder, Felix J, et al.
Published: (2024)
by: Binder, Felix J, et al.
Published: (2024)
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026)
by: Wu, Junfei, et al.
Published: (2026)
Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration
by: Lv, Hang, et al.
Published: (2026)
by: Lv, Hang, et al.
Published: (2026)
Zero-Overhead Introspection for Adaptive Test-Time Compute
by: Manvi, Rohin, et al.
Published: (2025)
by: Manvi, Rohin, et al.
Published: (2025)
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
by: AlShikh, Waseem, et al.
Published: (2025)
by: AlShikh, Waseem, et al.
Published: (2025)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
by: Liang, Kaiqu, et al.
Published: (2024)
by: Liang, Kaiqu, et al.
Published: (2024)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
by: Yang, Diji, et al.
Published: (2024)
by: Yang, Diji, et al.
Published: (2024)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)
by: Qu, Yuxiao, et al.
Published: (2024)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
LLM-Agnostic Semantic Representation Attack
by: Lian, Jiawei, et al.
Published: (2026)
by: Lian, Jiawei, et al.
Published: (2026)
PAFT: Prompt-Agnostic Fine-Tuning
by: Wei, Chenxing, et al.
Published: (2025)
by: Wei, Chenxing, et al.
Published: (2025)
Attacks by Content: Automated Fact-checking is an AI Security Issue
by: Schlichtkrull, Michael
Published: (2025)
by: Schlichtkrull, Michael
Published: (2025)
AI-AI Esthetic Collaboration with Explicit Semiotic Awareness and Emergent Grammar Development
by: Moldovan, Nicanor I.
Published: (2025)
by: Moldovan, Nicanor I.
Published: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
by: Li, YuQian, et al.
Published: (2025)
by: Li, YuQian, et al.
Published: (2025)
Unsupervised Translation of Emergent Communication
by: Levy, Ido, et al.
Published: (2025)
by: Levy, Ido, et al.
Published: (2025)
AI Content Self-Detection for Transformer-based Large Language Models
by: Caiado, Antônio Junior Alves, et al.
Published: (2023)
by: Caiado, Antônio Junior Alves, et al.
Published: (2023)
Reflect then Learn: Active Prompting for Information Extraction Guided by Introspective Confusion
by: Zhao, Dong, et al.
Published: (2025)
by: Zhao, Dong, et al.
Published: (2025)
Measuring Human Contribution in AI-Assisted Content Generation
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024)
by: Sprague, Zayne, et al.
Published: (2024)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
by: Zhu, Shuqi, et al.
Published: (2026)
by: Zhu, Shuqi, et al.
Published: (2026)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
Similar Items
-
Privileged Self-Access Matters for Introspection in AI
by: Song, Siyuan, et al.
Published: (2025) -
Language Models Fail to Introspect About Their Knowledge of Language
by: Song, Siyuan, et al.
Published: (2025) -
Are Language Models More Like Libraries or Like Librarians? Bibliotechnism, the Novel Reference Problem, and the Attitudes of LLMs
by: Lederman, Harvey, et al.
Published: (2024) -
The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models
by: Drucker, Daniel, et al.
Published: (2026) -
Emergent Introspective Awareness in Large Language Models
by: Lindsey, Jack
Published: (2026)