Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yalon, Noam Steinmetz, Goldstein, Ariel, Mudrik, Liad, Geva, Mor |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Estimating Knowledge in Large Language Models Without Generating a Single Token
par: Gottesman, Daniela, et autres
Publié: (2024)
par: Gottesman, Daniela, et autres
Publié: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
par: Yona, Gal, et autres
Publié: (2024)
par: Yona, Gal, et autres
Publié: (2024)
Detecting (Un)answerability in Large Language Models with Linear Directions
par: Lavi, Maor Juliet, et autres
Publié: (2025)
par: Lavi, Maor Juliet, et autres
Publié: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
par: Yang, Sohee, et autres
Publié: (2024)
par: Yang, Sohee, et autres
Publié: (2024)
Inferring Functionality of Attention Heads from their Parameters
par: Elhelo, Amit, et autres
Publié: (2024)
par: Elhelo, Amit, et autres
Publié: (2024)
Precise In-Parameter Concept Erasure in Large Language Models
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
par: Biran, Eden, et autres
Publié: (2024)
par: Biran, Eden, et autres
Publié: (2024)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
par: Yang, Sohee, et autres
Publié: (2024)
par: Yang, Sohee, et autres
Publié: (2024)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
par: Cohen, Ido, et autres
Publié: (2024)
par: Cohen, Ido, et autres
Publié: (2024)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
par: Yona, Itay, et autres
Publié: (2026)
par: Yona, Itay, et autres
Publié: (2026)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
par: Ivgi, Maor, et autres
Publié: (2024)
par: Ivgi, Maor, et autres
Publié: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
par: Shafran, Or, et autres
Publié: (2025)
par: Shafran, Or, et autres
Publié: (2025)
Hallucinations Undermine Trust; Metacognition is a Way Forward
par: Yona, Gal, et autres
Publié: (2026)
par: Yona, Gal, et autres
Publié: (2026)
Eliciting Textual Descriptions from Representations of Continuous Prompts
par: Ramati, Dana, et autres
Publié: (2024)
par: Ramati, Dana, et autres
Publié: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
par: Yona, Gal, et autres
Publié: (2024)
par: Yona, Gal, et autres
Publié: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
par: Katz, Shahar, et autres
Publié: (2024)
par: Katz, Shahar, et autres
Publié: (2024)
Motivation in Large Language Models
par: Nahum, Omer, et autres
Publié: (2026)
par: Nahum, Omer, et autres
Publié: (2026)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
par: Shafran, Or, et autres
Publié: (2026)
par: Shafran, Or, et autres
Publié: (2026)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
par: Huang, Jing, et autres
Publié: (2024)
par: Huang, Jing, et autres
Publié: (2024)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
par: Ahrac, Sagi, et autres
Publié: (2026)
par: Ahrac, Sagi, et autres
Publié: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
par: Avrahamy, Asaf, et autres
Publié: (2026)
par: Avrahamy, Asaf, et autres
Publié: (2026)
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
par: Goldstein, Ariel, et autres
Publié: (2024)
par: Goldstein, Ariel, et autres
Publié: (2024)
Preventing Rogue Agents Improves Multi-Agent Collaboration
par: Barbi, Ohav, et autres
Publié: (2025)
par: Barbi, Ohav, et autres
Publié: (2025)
The Hidden Space of Transformer Language Adapters
par: Alabi, Jesujoba O., et autres
Publié: (2024)
par: Alabi, Jesujoba O., et autres
Publié: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
par: Ghandeharioun, Asma, et autres
Publié: (2024)
par: Ghandeharioun, Asma, et autres
Publié: (2024)
Rethinking Selective Knowledge Distillation
par: Tavor, Almog, et autres
Publié: (2026)
par: Tavor, Almog, et autres
Publié: (2026)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
par: Din, Alexander Yom, et autres
Publié: (2023)
par: Din, Alexander Yom, et autres
Publié: (2023)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
par: Gottesman, Daniela, et autres
Publié: (2025)
par: Gottesman, Daniela, et autres
Publié: (2025)
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
par: Yang, Sohee, et autres
Publié: (2025)
par: Yang, Sohee, et autres
Publié: (2025)
Large Language Models are Capable of Offering Cognitive Reappraisal, if Guided
par: Zhan, Hongli, et autres
Publié: (2024)
par: Zhan, Hongli, et autres
Publié: (2024)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
par: Parmar, Mihir, et autres
Publié: (2022)
par: Parmar, Mihir, et autres
Publié: (2022)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
par: Mosbach, Marius, et autres
Publié: (2024)
par: Mosbach, Marius, et autres
Publié: (2024)
Deep Search with Hierarchical Meta-Cognitive Monitoring Inspired by Cognitive Neuroscience
par: Sun, Zhongxiang, et autres
Publié: (2026)
par: Sun, Zhongxiang, et autres
Publié: (2026)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
par: Hong, Yihuai, et autres
Publié: (2024)
par: Hong, Yihuai, et autres
Publié: (2024)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
par: Gekhman, Zorik, et autres
Publié: (2026)
par: Gekhman, Zorik, et autres
Publié: (2026)
Belief Revision: The Adaptability of Large Language Models Reasoning
par: Wilie, Bryan, et autres
Publié: (2024)
par: Wilie, Bryan, et autres
Publié: (2024)
HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model
par: Kayzer, Noam, et autres
Publié: (2026)
par: Kayzer, Noam, et autres
Publié: (2026)
Documents similaires
-
Estimating Knowledge in Large Language Models Without Generating a Single Token
par: Gottesman, Daniela, et autres
Publié: (2024) -
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
par: Yona, Gal, et autres
Publié: (2024) -
Detecting (Un)answerability in Large Language Models with Linear Directions
par: Lavi, Maor Juliet, et autres
Publié: (2025) -
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
par: Gur-Arieh, Yoav, et autres
Publié: (2026) -
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
par: Gur-Arieh, Yoav, et autres
Publié: (2025)