Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection
Fuente:
arXiv
Salvato in:
| Autori principali: | Naphade, Atharv, Bhargav, Samarth, Lim, Sean, Shah, Mcnair |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rational Synthesizers or Heuristic Followers? Analyzing LLMs in RAG-based Question-Answering
di: Naphade, Atharv
Pubblicazione: (2026)
di: Naphade, Atharv
Pubblicazione: (2026)
Interpreting Affine Recurrence Learning in GPT-style Transformers
di: Bhargav, Samarth, et al.
Pubblicazione: (2024)
di: Bhargav, Samarth, et al.
Pubblicazione: (2024)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
di: Schoene, Annika M, et al.
Pubblicazione: (2025)
di: Schoene, Annika M, et al.
Pubblicazione: (2025)
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
di: Lee, Jaeho, et al.
Pubblicazione: (2025)
di: Lee, Jaeho, et al.
Pubblicazione: (2025)
Introspective Diffusion Language Models
di: Yu, Yifan, et al.
Pubblicazione: (2026)
di: Yu, Yifan, et al.
Pubblicazione: (2026)
Introspection of Thought Helps AI Agents
di: Sun, Haoran, et al.
Pubblicazione: (2025)
di: Sun, Haoran, et al.
Pubblicazione: (2025)
"I May Not Have Articulated Myself Clearly": Diagnosing Dynamic Instability in LLM Reasoning at Inference Time
di: Chen, Jinkun, et al.
Pubblicazione: (2026)
di: Chen, Jinkun, et al.
Pubblicazione: (2026)
Can LLMs Introspect? A Reality Check
di: Singh, Shashwat, et al.
Pubblicazione: (2026)
di: Singh, Shashwat, et al.
Pubblicazione: (2026)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
di: Shenoy, Keshav, et al.
Pubblicazione: (2026)
di: Shenoy, Keshav, et al.
Pubblicazione: (2026)
Emergent Introspection in AI is Content-Agnostic
di: Lederman, Harvey, et al.
Pubblicazione: (2026)
di: Lederman, Harvey, et al.
Pubblicazione: (2026)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
di: Cui, Brandon, et al.
Pubblicazione: (2026)
di: Cui, Brandon, et al.
Pubblicazione: (2026)
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
di: Hahami, Ely, et al.
Pubblicazione: (2025)
di: Hahami, Ely, et al.
Pubblicazione: (2025)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
di: Ranganath, Suraj, et al.
Pubblicazione: (2026)
di: Ranganath, Suraj, et al.
Pubblicazione: (2026)
Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop
di: Goel, Atharv, et al.
Pubblicazione: (2025)
di: Goel, Atharv, et al.
Pubblicazione: (2025)
Emergent Introspective Awareness in Large Language Models
di: Lindsey, Jack
Pubblicazione: (2026)
di: Lindsey, Jack
Pubblicazione: (2026)
Privileged Self-Access Matters for Introspection in AI
di: Song, Siyuan, et al.
Pubblicazione: (2025)
di: Song, Siyuan, et al.
Pubblicazione: (2025)
InnerPond: Fostering Inter-Self Dialogue with a Multi-Agent Approach for Introspection
di: Jeon, Hayeon, et al.
Pubblicazione: (2026)
di: Jeon, Hayeon, et al.
Pubblicazione: (2026)
Agency in the Age of AI
di: Swarup, Samarth
Pubblicazione: (2025)
di: Swarup, Samarth
Pubblicazione: (2025)
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
di: Martorell, Nicolas, et al.
Pubblicazione: (2026)
di: Martorell, Nicolas, et al.
Pubblicazione: (2026)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
Language Models Fail to Introspect About Their Knowledge of Language
di: Song, Siyuan, et al.
Pubblicazione: (2025)
di: Song, Siyuan, et al.
Pubblicazione: (2025)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
di: Mishra, Samarth, et al.
Pubblicazione: (2025)
di: Mishra, Samarth, et al.
Pubblicazione: (2025)
Position: Introspective Experience from Conversational Environments as a Path to Better Learning
di: Musat, Claudiu Cristian, et al.
Pubblicazione: (2026)
di: Musat, Claudiu Cristian, et al.
Pubblicazione: (2026)
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
Systems Explaining Systems: A Framework for Intelligence and Consciousness
di: Semmler, Sean Niklas
Pubblicazione: (2026)
di: Semmler, Sean Niklas
Pubblicazione: (2026)
Towards Optimizing the Costs of LLM Usage
di: Shekhar, Shivanshu, et al.
Pubblicazione: (2024)
di: Shekhar, Shivanshu, et al.
Pubblicazione: (2024)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
di: Veldanda, Akshaj Kumar, et al.
Pubblicazione: (2024)
di: Veldanda, Akshaj Kumar, et al.
Pubblicazione: (2024)
Context Branching for LLM Conversations: A Version Control Approach to Exploratory Programming
di: Nanjundappa, Bhargav Chickmagalur, et al.
Pubblicazione: (2025)
di: Nanjundappa, Bhargav Chickmagalur, et al.
Pubblicazione: (2025)
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
di: Su, Guinan, et al.
Pubblicazione: (2025)
di: Su, Guinan, et al.
Pubblicazione: (2025)
Exploration Through Introspection: A Self-Aware Reward Model
di: Petrowski, Michael, et al.
Pubblicazione: (2026)
di: Petrowski, Michael, et al.
Pubblicazione: (2026)
Latent Introspection: Models Can Detect Prior Concept Injections
di: Pearson-Vogel, Theia, et al.
Pubblicazione: (2026)
di: Pearson-Vogel, Theia, et al.
Pubblicazione: (2026)
Does It Make Sense to Speak of Introspection in Large Language Models?
di: Comsa, Iulia M., et al.
Pubblicazione: (2025)
di: Comsa, Iulia M., et al.
Pubblicazione: (2025)
Chatting with Images for Introspective Visual Thinking
di: Wu, Junfei, et al.
Pubblicazione: (2026)
di: Wu, Junfei, et al.
Pubblicazione: (2026)
AI, Help Me Think$\unicode{x2014}$but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support
di: Reicherts, Leon, et al.
Pubblicazione: (2025)
di: Reicherts, Leon, et al.
Pubblicazione: (2025)
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
di: He, Zeyu, et al.
Pubblicazione: (2025)
di: He, Zeyu, et al.
Pubblicazione: (2025)
Looking Inward: Language Models Can Learn About Themselves by Introspection
di: Binder, Felix J, et al.
Pubblicazione: (2024)
di: Binder, Felix J, et al.
Pubblicazione: (2024)
Zero-Overhead Introspection for Adaptive Test-Time Compute
di: Manvi, Rohin, et al.
Pubblicazione: (2025)
di: Manvi, Rohin, et al.
Pubblicazione: (2025)
Distributive Fairness in Large Language Models: Evaluating Alignment with Human Values
di: Hosseini, Hadi, et al.
Pubblicazione: (2025)
di: Hosseini, Hadi, et al.
Pubblicazione: (2025)
The Unlearning Mirage: A Dynamic Framework for Evaluating LLM Unlearning
di: Shah, Raj Sanjay, et al.
Pubblicazione: (2026)
di: Shah, Raj Sanjay, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Rational Synthesizers or Heuristic Followers? Analyzing LLMs in RAG-based Question-Answering
di: Naphade, Atharv
Pubblicazione: (2026) -
Interpreting Affine Recurrence Learning in GPT-style Transformers
di: Bhargav, Samarth, et al.
Pubblicazione: (2024) -
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
di: Laine, Rudolf, et al.
Pubblicazione: (2024) -
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
di: Schoene, Annika M, et al.
Pubblicazione: (2025) -
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
di: Lee, Jaeho, et al.
Pubblicazione: (2025)