Standards for Belief Representations in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Herrmann, Daniel A., Levinstein, Benjamin A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Does ChatGPT Have a Mind?
von: Goldstein, Simon, et al.
Veröffentlicht: (2024)
von: Goldstein, Simon, et al.
Veröffentlicht: (2024)
A Decision-Theoretic Approach for Managing Misalignment
von: Herrmann, Daniel A., et al.
Veröffentlicht: (2025)
von: Herrmann, Daniel A., et al.
Veröffentlicht: (2025)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
von: Huang, Fan, et al.
Veröffentlicht: (2026)
von: Huang, Fan, et al.
Veröffentlicht: (2026)
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
von: Herrmann, Nils A., et al.
Veröffentlicht: (2026)
von: Herrmann, Nils A., et al.
Veröffentlicht: (2026)
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
von: Levinstein, B. A., et al.
Veröffentlicht: (2023)
von: Levinstein, B. A., et al.
Veröffentlicht: (2023)
Probing the Lack of Stable Internal Beliefs in LLMs
von: Luo, Yifan, et al.
Veröffentlicht: (2026)
von: Luo, Yifan, et al.
Veröffentlicht: (2026)
Using LLMs to Model the Beliefs and Preferences of Targeted Populations
von: Namikoshi, Keiichi, et al.
Veröffentlicht: (2024)
von: Namikoshi, Keiichi, et al.
Veröffentlicht: (2024)
Mind the (Belief) Gap: Group Identity in the World of LLMs
von: Borah, Angana, et al.
Veröffentlicht: (2025)
von: Borah, Angana, et al.
Veröffentlicht: (2025)
Collaborative Belief Reasoning with LLMs for Efficient Multi-Agent Collaboration
von: Wang, Zhimin, et al.
Veröffentlicht: (2025)
von: Wang, Zhimin, et al.
Veröffentlicht: (2025)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Belief-State Query Policies for User-Aligned POMDPs
von: Bramblett, Daniel, et al.
Veröffentlicht: (2024)
von: Bramblett, Daniel, et al.
Veröffentlicht: (2024)
When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
von: Yamin, Khurram, et al.
Veröffentlicht: (2026)
von: Yamin, Khurram, et al.
Veröffentlicht: (2026)
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
von: Liu, Xuelin, et al.
Veröffentlicht: (2024)
von: Liu, Xuelin, et al.
Veröffentlicht: (2024)
Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
von: Bortoletto, Matteo, et al.
Veröffentlicht: (2024)
von: Bortoletto, Matteo, et al.
Veröffentlicht: (2024)
Can LLMs Emulate Human Belief Dynamics?
von: Proma, Adiba Mahbub, et al.
Veröffentlicht: (2026)
von: Proma, Adiba Mahbub, et al.
Veröffentlicht: (2026)
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
von: Marklund, Henrik, et al.
Veröffentlicht: (2024)
von: Marklund, Henrik, et al.
Veröffentlicht: (2024)
An Abstract Worlds Semantic Framework for Belief Change Operators
von: Grimaldi, Daniel, et al.
Veröffentlicht: (2026)
von: Grimaldi, Daniel, et al.
Veröffentlicht: (2026)
Common Belief Revisited
von: Ågotnes, Thomas
Veröffentlicht: (2026)
von: Ågotnes, Thomas
Veröffentlicht: (2026)
Learning Dynamic Belief Graphs for Theory-of-mind Reasoning
von: Chen, Ruxiao, et al.
Veröffentlicht: (2026)
von: Chen, Ruxiao, et al.
Veröffentlicht: (2026)
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
von: Hase, Peter, et al.
Veröffentlicht: (2024)
von: Hase, Peter, et al.
Veröffentlicht: (2024)
LLMs as Strategic Agents: Beliefs, Best Response Behavior, and Emergent Heuristics
von: de Fortuny, Enric Junque, et al.
Veröffentlicht: (2025)
von: de Fortuny, Enric Junque, et al.
Veröffentlicht: (2025)
How Do People Revise Inconsistent Beliefs? Examining Belief Revision in Humans with User Studies
von: Vasileiou, Stylianos Loukas, et al.
Veröffentlicht: (2025)
von: Vasileiou, Stylianos Loukas, et al.
Veröffentlicht: (2025)
FairBelief -- Assessing Harmful Beliefs in Language Models
von: Setzu, Mattia, et al.
Veröffentlicht: (2024)
von: Setzu, Mattia, et al.
Veröffentlicht: (2024)
Deep Belief Markov Models for POMDP Inference
von: Arcieri, Giacomo, et al.
Veröffentlicht: (2025)
von: Arcieri, Giacomo, et al.
Veröffentlicht: (2025)
Localizing Persona Representations in LLMs
von: Cintas, Celia, et al.
Veröffentlicht: (2025)
von: Cintas, Celia, et al.
Veröffentlicht: (2025)
Uncovering the Computational Ingredients of Human-Like Representations in LLMs
von: Studdiford, Zach, et al.
Veröffentlicht: (2025)
von: Studdiford, Zach, et al.
Veröffentlicht: (2025)
On Definite Iterated Belief Revision with Belief Algebras
von: Meng, Hua, et al.
Veröffentlicht: (2025)
von: Meng, Hua, et al.
Veröffentlicht: (2025)
ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs
von: Zhang, Yunuo, et al.
Veröffentlicht: (2025)
von: Zhang, Yunuo, et al.
Veröffentlicht: (2025)
Belief Change based on Knowledge Measures
von: Straccia, Umberto, et al.
Veröffentlicht: (2024)
von: Straccia, Umberto, et al.
Veröffentlicht: (2024)
Are LLMs Socially Adaptive? Contrasting Belief Evolution in Large Language Models and Humans
von: Lei, Yu, et al.
Veröffentlicht: (2024)
von: Lei, Yu, et al.
Veröffentlicht: (2024)
Belief or Circuitry? Causal Evidence for In-Context Graph Learning
von: Kowalyshyn, Katharine, et al.
Veröffentlicht: (2026)
von: Kowalyshyn, Katharine, et al.
Veröffentlicht: (2026)
Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs
von: Liu, Xuyuan, et al.
Veröffentlicht: (2025)
von: Liu, Xuyuan, et al.
Veröffentlicht: (2025)
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
von: Borah, Angana, et al.
Veröffentlicht: (2026)
von: Borah, Angana, et al.
Veröffentlicht: (2026)
Uncommon Belief in Rationality
von: Shi, Qi, et al.
Veröffentlicht: (2024)
von: Shi, Qi, et al.
Veröffentlicht: (2024)
Belief sharing: a blessing or a curse
von: Catal, Ozan, et al.
Veröffentlicht: (2024)
von: Catal, Ozan, et al.
Veröffentlicht: (2024)
Recursive Belief Vision Language Action Models
von: Bagaria, Vaidehi, et al.
Veröffentlicht: (2026)
von: Bagaria, Vaidehi, et al.
Veröffentlicht: (2026)
Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial Observability
von: Kuo, Po-Chen, et al.
Veröffentlicht: (2025)
von: Kuo, Po-Chen, et al.
Veröffentlicht: (2025)
Where Common Knowledge Cannot Be Formed, Common Belief Can -- Planning with Multi-Agent Belief Using Group Justified Perspectives
von: Hu, Guang, et al.
Veröffentlicht: (2024)
von: Hu, Guang, et al.
Veröffentlicht: (2024)
Reasoning Effort and Problem Complexity: A Scaling Analysis in LLMs
von: Estermann, Benjamin, et al.
Veröffentlicht: (2025)
von: Estermann, Benjamin, et al.
Veröffentlicht: (2025)
Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems
von: Kostka, Adam, et al.
Veröffentlicht: (2026)
von: Kostka, Adam, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Does ChatGPT Have a Mind?
von: Goldstein, Simon, et al.
Veröffentlicht: (2024) -
A Decision-Theoretic Approach for Managing Misalignment
von: Herrmann, Daniel A., et al.
Veröffentlicht: (2025) -
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
von: Huang, Fan, et al.
Veröffentlicht: (2026) -
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
von: Herrmann, Nils A., et al.
Veröffentlicht: (2026) -
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
von: Levinstein, B. A., et al.
Veröffentlicht: (2023)