Probing the Lack of Stable Internal Beliefs in LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Luo, Yifan, Xu, Kangping, Lu, Yanzhen, Yuan, Yang, Yao, Andrew Chi-Chih |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Towards Automated Formal Verification of Backend Systems with LLMs
par: Xu, Kangping, et autres
Publié: (2025)
par: Xu, Kangping, et autres
Publié: (2025)
Meta Prompting for AI Systems
par: Zhang, Yifan, et autres
Publié: (2023)
par: Zhang, Yifan, et autres
Publié: (2023)
Group Representational Position Encoding
par: Zhang, Yifan, et autres
Publié: (2025)
par: Zhang, Yifan, et autres
Publié: (2025)
On the Diagram of Thought
par: Zhang, Yifan, et autres
Publié: (2024)
par: Zhang, Yifan, et autres
Publié: (2024)
Monadic Context Engineering
par: Zhang, Yifan, et autres
Publié: (2025)
par: Zhang, Yifan, et autres
Publié: (2025)
Augmenting Math Word Problems via Iterative Question Composing
par: Liu, Haoxiong, et autres
Publié: (2024)
par: Liu, Haoxiong, et autres
Publié: (2024)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
par: Zhang, Yifan, et autres
Publié: (2025)
par: Zhang, Yifan, et autres
Publié: (2025)
Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
par: Zhang, Yifan, et autres
Publié: (2024)
par: Zhang, Yifan, et autres
Publié: (2024)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
par: Zhang, Yifan, et autres
Publié: (2024)
par: Zhang, Yifan, et autres
Publié: (2024)
Tensor Product Attention Is All You Need
par: Zhang, Yifan, et autres
Publié: (2025)
par: Zhang, Yifan, et autres
Publié: (2025)
Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations
par: Chi, Jinwei, et autres
Publié: (2025)
par: Chi, Jinwei, et autres
Publié: (2025)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
par: Chang, Chen-Chi, et autres
Publié: (2024)
par: Chang, Chen-Chi, et autres
Publié: (2024)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
par: Huang, Fan, et autres
Publié: (2026)
par: Huang, Fan, et autres
Publié: (2026)
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
par: Ezra, Elon, et autres
Publié: (2025)
par: Ezra, Elon, et autres
Publié: (2025)
From Human Cognition to Neural Activations: Probing the Computational Primitives of Spatial Reasoning in LLMs
par: An, Jiyuan, et autres
Publié: (2026)
par: An, Jiyuan, et autres
Publié: (2026)
Hallucination Detection with the Internal Layers of LLMs
par: Preiß, Martin
Publié: (2025)
par: Preiß, Martin
Publié: (2025)
Mind the (Belief) Gap: Group Identity in the World of LLMs
par: Borah, Angana, et autres
Publié: (2025)
par: Borah, Angana, et autres
Publié: (2025)
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
par: Chen, Jiuting, et autres
Publié: (2026)
par: Chen, Jiuting, et autres
Publié: (2026)
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
par: Wang, Yanli, et autres
Publié: (2026)
par: Wang, Yanli, et autres
Publié: (2026)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
par: Agarwal, Dhruv, et autres
Publié: (2025)
par: Agarwal, Dhruv, et autres
Publié: (2025)
Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty
par: Singh, Joykirat, et autres
Publié: (2026)
par: Singh, Joykirat, et autres
Publié: (2026)
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
par: Liu, Xuelin, et autres
Publié: (2024)
par: Liu, Xuelin, et autres
Publié: (2024)
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
par: Clymer, Joshua, et autres
Publié: (2024)
par: Clymer, Joshua, et autres
Publié: (2024)
MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
par: Dong, Haoyu, et autres
Publié: (2025)
par: Dong, Haoyu, et autres
Publié: (2025)
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
par: Luo, Yifan, et autres
Publié: (2024)
par: Luo, Yifan, et autres
Publié: (2024)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
par: Ni, Jingwei, et autres
Publié: (2025)
par: Ni, Jingwei, et autres
Publié: (2025)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
par: Xu, Rongwu, et autres
Publié: (2023)
par: Xu, Rongwu, et autres
Publié: (2023)
Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
par: Lee, Seongmin, et autres
Publié: (2024)
par: Lee, Seongmin, et autres
Publié: (2024)
Finding the Cracks: Improving LLMs Reasoning with Paraphrastic Probing and Consistency Verification
par: Shi, Weili, et autres
Publié: (2026)
par: Shi, Weili, et autres
Publié: (2026)
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
par: Hase, Peter, et autres
Publié: (2024)
par: Hase, Peter, et autres
Publié: (2024)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
par: Chiu, Yu Ying, et autres
Publié: (2024)
par: Chiu, Yu Ying, et autres
Publié: (2024)
HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
par: Liu, Jiaxin, et autres
Publié: (2025)
par: Liu, Jiaxin, et autres
Publié: (2025)
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
par: Kale, Sahil
Publié: (2025)
par: Kale, Sahil
Publié: (2025)
Beyond Words: A Latent Memory Approach to Internal Reasoning in LLMs
par: Orlicki, José I.
Publié: (2025)
par: Orlicki, José I.
Publié: (2025)
Agentic Recommender System with Hierarchical Belief-State Memory
par: Shen, Xiang, et autres
Publié: (2026)
par: Shen, Xiang, et autres
Publié: (2026)
FairBelief -- Assessing Harmful Beliefs in Language Models
par: Setzu, Mattia, et autres
Publié: (2024)
par: Setzu, Mattia, et autres
Publié: (2024)
Using LLMs to Model the Beliefs and Preferences of Targeted Populations
par: Namikoshi, Keiichi, et autres
Publié: (2024)
par: Namikoshi, Keiichi, et autres
Publié: (2024)
TECP: Token-Entropy Conformal Prediction for LLMs
par: Xu, Beining, et autres
Publié: (2025)
par: Xu, Beining, et autres
Publié: (2025)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
par: Chen, Justin Chih-Yao, et autres
Publié: (2023)
par: Chen, Justin Chih-Yao, et autres
Publié: (2023)
Are Your LLMs Capable of Stable Reasoning?
par: Liu, Junnan, et autres
Publié: (2024)
par: Liu, Junnan, et autres
Publié: (2024)
Documents similaires
-
Towards Automated Formal Verification of Backend Systems with LLMs
par: Xu, Kangping, et autres
Publié: (2025) -
Meta Prompting for AI Systems
par: Zhang, Yifan, et autres
Publié: (2023) -
Group Representational Position Encoding
par: Zhang, Yifan, et autres
Publié: (2025) -
On the Diagram of Thought
par: Zhang, Yifan, et autres
Publié: (2024) -
Monadic Context Engineering
par: Zhang, Yifan, et autres
Publié: (2025)