Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hahami, Ely, Sinha, Ishaan, Jain, Lavik, Kaplan, Josh, Hahami, Jon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memorization and Knowledge Injection in Gated LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
Can LLMs Introspect? A Reality Check
by: Singh, Shashwat, et al.
Published: (2026)
by: Singh, Shashwat, et al.
Published: (2026)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
by: Shenoy, Keshav, et al.
Published: (2026)
by: Shenoy, Keshav, et al.
Published: (2026)
Toward Revealing Nuanced Biases in Medical LLMs
by: Adiba, Farzana Islam, et al.
Published: (2025)
by: Adiba, Farzana Islam, et al.
Published: (2025)
Metacognition is all you need? Using Introspection in Generative Agents to Improve Goal-directed Behavior
by: Toy, Jason, et al.
Published: (2024)
by: Toy, Jason, et al.
Published: (2024)
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Introspective Diffusion Language Models
by: Yu, Yifan, et al.
Published: (2026)
by: Yu, Yifan, et al.
Published: (2026)
Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization
by: Verma, Ishaan, et al.
Published: (2025)
by: Verma, Ishaan, et al.
Published: (2025)
Introspection of Thought Helps AI Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Latent Introspection: Models Can Detect Prior Concept Injections
by: Pearson-Vogel, Theia, et al.
Published: (2026)
by: Pearson-Vogel, Theia, et al.
Published: (2026)
NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
by: Carvalho, Wilka, et al.
Published: (2025)
by: Carvalho, Wilka, et al.
Published: (2025)
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances
by: Wu, Zehui, et al.
Published: (2024)
by: Wu, Zehui, et al.
Published: (2024)
Unified Smart Factory Model: A model-based Approach for Integrating Industry 4.0 and Sustainability for Manufacturing Systems
by: Kaushal, Ishaan, et al.
Published: (2025)
by: Kaushal, Ishaan, et al.
Published: (2025)
Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams
by: Bansal, Aayam, et al.
Published: (2025)
by: Bansal, Aayam, et al.
Published: (2025)
Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection
by: Naphade, Atharv, et al.
Published: (2026)
by: Naphade, Atharv, et al.
Published: (2026)
Emergent Introspection in AI is Content-Agnostic
by: Lederman, Harvey, et al.
Published: (2026)
by: Lederman, Harvey, et al.
Published: (2026)
MicroProbe: Efficient Reliability Assessment for Foundation Models with Minimal Data
by: Bansal, Aayam, et al.
Published: (2025)
by: Bansal, Aayam, et al.
Published: (2025)
AgentComm-Bench: Stress-Testing Cooperative Embodied AI Under Latency, Packet Loss, and Bandwidth Collapse
by: Bansal, Aayam, et al.
Published: (2026)
by: Bansal, Aayam, et al.
Published: (2026)
Privileged Self-Access Matters for Introspection in AI
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Emergent Introspective Awareness in Large Language Models
by: Lindsey, Jack
Published: (2026)
by: Lindsey, Jack
Published: (2026)
Enhancing Low-Resource Minority Language Translation with LLMs and Retrieval-Augmented Generation for Cultural Nuances
by: Chang, Chen-Chi, et al.
Published: (2025)
by: Chang, Chen-Chi, et al.
Published: (2025)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
by: Domkundwar, Ishaan, et al.
Published: (2024)
by: Domkundwar, Ishaan, et al.
Published: (2024)
CAND: Cross-Domain Ambiguity Inference for Early Detecting Nuanced Illness Deterioration
by: Ting, Lo Pang-Yun, et al.
Published: (2025)
by: Ting, Lo Pang-Yun, et al.
Published: (2025)
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
Improving LLMs' Generalized Reasoning Abilities by Graph Problems
by: Zhang, Qifan, et al.
Published: (2025)
by: Zhang, Qifan, et al.
Published: (2025)
Exploration Through Introspection: A Self-Aware Reward Model
by: Petrowski, Michael, et al.
Published: (2026)
by: Petrowski, Michael, et al.
Published: (2026)
Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs
by: Cheng, Kewei, et al.
Published: (2024)
by: Cheng, Kewei, et al.
Published: (2024)
Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
by: Bhola, Ishaan, et al.
Published: (2025)
by: Bhola, Ishaan, et al.
Published: (2025)
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
by: Martorell, Nicolas, et al.
Published: (2026)
by: Martorell, Nicolas, et al.
Published: (2026)
Language Models Fail to Introspect About Their Knowledge of Language
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
The Counterfeit Conundrum: Can Code Language Models Grasp the Nuances of Their Incorrect Generations?
by: Gu, Alex, et al.
Published: (2024)
by: Gu, Alex, et al.
Published: (2024)
Longitudinal Boundary Sharpness Coefficient Slopes Predict Time to Alzheimer's Disease Conversion in Mild Cognitive Impairment: A Survival Analysis Using the ADNI Cohort
by: Cherukuri, Ishaan
Published: (2026)
by: Cherukuri, Ishaan
Published: (2026)
Position: Introspective Experience from Conversational Environments as a Path to Better Learning
by: Musat, Claudiu Cristian, et al.
Published: (2026)
by: Musat, Claudiu Cristian, et al.
Published: (2026)
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)
by: Naik, Ranjita, et al.
Published: (2023)
SuperCoder2.0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer
by: Gautam, Anmol, et al.
Published: (2024)
by: Gautam, Anmol, et al.
Published: (2024)
Does It Make Sense to Speak of Introspection in Large Language Models?
by: Comsa, Iulia M., et al.
Published: (2025)
by: Comsa, Iulia M., et al.
Published: (2025)
Light-Weight Benchmarks Reveal the Hidden Hardware Cost of Zero-Shot Tabular Foundation Models
by: Gangwani, Ishaan, et al.
Published: (2025)
by: Gangwani, Ishaan, et al.
Published: (2025)
Reason-to-Transmit: Deliberative Adaptive Communication for Cooperative Perception
by: Bansal, Aayam, et al.
Published: (2026)
by: Bansal, Aayam, et al.
Published: (2026)
Similar Items
-
Memorization and Knowledge Injection in Gated LLMs
by: Pan, Xu, et al.
Published: (2025) -
Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
by: Pan, Xu, et al.
Published: (2025) -
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025) -
Can LLMs Introspect? A Reality Check
by: Singh, Shashwat, et al.
Published: (2026) -
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
by: Shenoy, Keshav, et al.
Published: (2026)