Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs
Fuente:
arXiv
Saved in:
| Main Authors: | Trott, Sean, Taylor, Samuel, Jones, Cameron, Michaelov, James A., Rivière, Pamela D. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Open Must Language Models be to Enable Reliable Scientific Inference?
by: Michaelov, James A., et al.
Published: (2026)
by: Michaelov, James A., et al.
Published: (2026)
Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
by: Trott, Sean
Published: (2025)
by: Trott, Sean
Published: (2025)
Capacity Constraints and the Multilingual Penalty for Lexical Disambiguation
by: Trott, Sean, et al.
Published: (2026)
by: Trott, Sean, et al.
Published: (2026)
Measuring and Modifying the Readability of English Texts with GPT-4
by: Trott, Sean, et al.
Published: (2024)
by: Trott, Sean, et al.
Published: (2024)
Start Making Sense(s): A Developmental Probe of Attention Specialization Using Lexical Ambiguity
by: Rivière, Pamela D., et al.
Published: (2025)
by: Rivière, Pamela D., et al.
Published: (2025)
Do language models capture implied discourse meanings? An investigation with exhaustivity implicatures of Korean morphology
by: Shin, Hagyeong, et al.
Published: (2024)
by: Shin, Hagyeong, et al.
Published: (2024)
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
by: Michaelov, James A., et al.
Published: (2025)
by: Michaelov, James A., et al.
Published: (2025)
Great Memory, Shallow Reasoning: Limits of $k$NN-LMs
by: Geng, Shangyi, et al.
Published: (2024)
by: Geng, Shangyi, et al.
Published: (2024)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
by: He, Zoe Wanying, et al.
Published: (2025)
by: He, Zoe Wanying, et al.
Published: (2025)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
by: Damani, Mehul, et al.
Published: (2025)
by: Damani, Mehul, et al.
Published: (2025)
Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models
by: Bose, Kartik, et al.
Published: (2026)
by: Bose, Kartik, et al.
Published: (2026)
Evaluating Contextualized Representations of (Spanish) Ambiguous Words: A New Lexical Resource and Empirical Analysis
by: Rivière, Pamela D., et al.
Published: (2024)
by: Rivière, Pamela D., et al.
Published: (2024)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning
by: Zhang, Yadong, et al.
Published: (2024)
by: Zhang, Yadong, et al.
Published: (2024)
Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon
by: Prashanth, USVSN Sai, et al.
Published: (2024)
by: Prashanth, USVSN Sai, et al.
Published: (2024)
FairBelief -- Assessing Harmful Beliefs in Language Models
by: Setzu, Mattia, et al.
Published: (2024)
by: Setzu, Mattia, et al.
Published: (2024)
Examining False Positives under Inference Scaling for Mathematical Reasoning
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
by: Asai, Akari, et al.
Published: (2024)
by: Asai, Akari, et al.
Published: (2024)
A Statistical Physics of Language Model Reasoning
by: Carson, Jack David, et al.
Published: (2025)
by: Carson, Jack David, et al.
Published: (2025)
Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
by: Bochkov, A.
Published: (2025)
by: Bochkov, A.
Published: (2025)
Where Should I Study? Biased Language Models Decide! Evaluating Fairness in LMs for Academic Recommendations
by: Shailya, Krithi, et al.
Published: (2025)
by: Shailya, Krithi, et al.
Published: (2025)
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking
by: Zhu, Wang Bill, et al.
Published: (2026)
by: Zhu, Wang Bill, et al.
Published: (2026)
Zero, Finite, and Infinite Belief History of Theory of Mind Reasoning in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)
by: Tang, Weizhi, et al.
Published: (2024)
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models
by: Kaiser, Caspar, et al.
Published: (2026)
by: Kaiser, Caspar, et al.
Published: (2026)
ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization
by: Ahuja, Riyaz, et al.
Published: (2026)
by: Ahuja, Riyaz, et al.
Published: (2026)
Localizing AI: Evaluating Open-Weight Language Models for Languages of Baltic States
by: Kapočiūtė-Dzikienė, Jurgita, et al.
Published: (2025)
by: Kapočiūtė-Dzikienė, Jurgita, et al.
Published: (2025)
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
by: Chatterjee, Sagnik, et al.
Published: (2026)
by: Chatterjee, Sagnik, et al.
Published: (2026)
What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models
by: Zhu, Shengqi, et al.
Published: (2024)
by: Zhu, Shengqi, et al.
Published: (2024)
Selecting Language Models for Social Science: Start Small, Start Open, and Validate
by: Stoltz, Dustin S., et al.
Published: (2026)
by: Stoltz, Dustin S., et al.
Published: (2026)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
by: Hyeon, Sieun, et al.
Published: (2024)
by: Hyeon, Sieun, et al.
Published: (2024)
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
by: Leong, Chak Tou, et al.
Published: (2026)
by: Leong, Chak Tou, et al.
Published: (2026)
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
by: Clark, Peter, et al.
Published: (2023)
by: Clark, Peter, et al.
Published: (2023)
Lost in Translation: The Algorithmic Gap Between LMs and the Brain
by: Tosato, Tommaso, et al.
Published: (2024)
by: Tosato, Tommaso, et al.
Published: (2024)
DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs
by: Yu, Longxuan, et al.
Published: (2026)
by: Yu, Longxuan, et al.
Published: (2026)
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity
by: Burnat, Florian A. D., et al.
Published: (2026)
by: Burnat, Florian A. D., et al.
Published: (2026)
Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
by: Yoon, Hong-Jun, et al.
Published: (2025)
by: Yoon, Hong-Jun, et al.
Published: (2025)
Accumulating Context Changes the Beliefs of Language Models
by: Geng, Jiayi, et al.
Published: (2025)
by: Geng, Jiayi, et al.
Published: (2025)
Similar Items
-
How Open Must Language Models be to Enable Reliable Scientific Inference?
by: Michaelov, James A., et al.
Published: (2026) -
Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
by: Trott, Sean
Published: (2025) -
Capacity Constraints and the Multilingual Penalty for Lexical Disambiguation
by: Trott, Sean, et al.
Published: (2026) -
Measuring and Modifying the Readability of English Texts with GPT-4
by: Trott, Sean, et al.
Published: (2024) -
Start Making Sense(s): A Developmental Probe of Attention Specialization Using Lexical Ambiguity
by: Rivière, Pamela D., et al.
Published: (2025)