What Does it Mean for a Neural Network to Learn a "World Model"?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Kenneth, Viégas, Fernanda, Wattenberg, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
by: Li, Kenneth, et al.
Published: (2022)
by: Li, Kenneth, et al.
Published: (2022)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
by: Li, Kenneth, et al.
Published: (2023)
by: Li, Kenneth, et al.
Published: (2023)
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Relational Composition in Neural Networks: A Survey and Call to Action
by: Wattenberg, Martin, et al.
Published: (2024)
by: Wattenberg, Martin, et al.
Published: (2024)
Designing a Dashboard for Transparency and Control of Conversational AI
by: Chen, Yida, et al.
Published: (2024)
by: Chen, Yida, et al.
Published: (2024)
What Generative Artificial Intelligence Means for Terminological Definitions
by: Martín, Antonio San
Published: (2024)
by: Martín, Antonio San
Published: (2024)
What is Sentiment Meant to Mean to Language Models?
by: Burnham, Michael
Published: (2024)
by: Burnham, Michael
Published: (2024)
Shared Global and Local Geometry of Language Model Embeddings
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
by: Banerjee, Mohor, et al.
Published: (2025)
by: Banerjee, Mohor, et al.
Published: (2025)
Can Interpretation Predict Behavior on Unseen Data?
by: Li, Victoria R., et al.
Published: (2025)
by: Li, Victoria R., et al.
Published: (2025)
Does visualization help AI understand data?
by: Li, Victoria R., et al.
Published: (2025)
by: Li, Victoria R., et al.
Published: (2025)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
by: Verdini, Francesco, et al.
Published: (2024)
by: Verdini, Francesco, et al.
Published: (2024)
ICLR: In-Context Learning of Representations
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Assessing the Capability of LLMs in Solving POSCOMP Questions
by: Viegas, Cayo, et al.
Published: (2025)
by: Viegas, Cayo, et al.
Published: (2025)
Chronotome: Real-Time Topic Modeling for Streaming Embedding Spaces
by: Lim, Matte, et al.
Published: (2025)
by: Lim, Matte, et al.
Published: (2025)
Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
by: Hu, Mengkang, et al.
Published: (2025)
by: Hu, Mengkang, et al.
Published: (2025)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
by: Orgad, Hadas, et al.
Published: (2026)
by: Orgad, Hadas, et al.
Published: (2026)
What is it for a Machine Learning Model to Have a Capability?
by: Harding, Jacqueline, et al.
Published: (2024)
by: Harding, Jacqueline, et al.
Published: (2024)
Word Meanings in Transformer Language Models
by: Grindrod, Jumbly, et al.
Published: (2025)
by: Grindrod, Jumbly, et al.
Published: (2025)
Evaluating the World Model Implicit in a Generative Model
by: Vafa, Keyon, et al.
Published: (2024)
by: Vafa, Keyon, et al.
Published: (2024)
Enhancing Agent Learning through World Dynamics Modeling
by: Sun, Zhiyuan, et al.
Published: (2024)
by: Sun, Zhiyuan, et al.
Published: (2024)
What is a Number, That a Large Language Model May Know It?
by: Marjieh, Raja, et al.
Published: (2025)
by: Marjieh, Raja, et al.
Published: (2025)
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
by: Blum, Carter, et al.
Published: (2025)
by: Blum, Carter, et al.
Published: (2025)
PatchWorld: Gradient-Free Optimization of Executable World Models
by: Bai, Jiaxin, et al.
Published: (2026)
by: Bai, Jiaxin, et al.
Published: (2026)
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
by: Hu, Mengkang, et al.
Published: (2025)
by: Hu, Mengkang, et al.
Published: (2025)
Does ChatGPT Have a Mind?
by: Goldstein, Simon, et al.
Published: (2024)
by: Goldstein, Simon, et al.
Published: (2024)
How to Learn in a Noisy World? Self-Correcting the Real-World Data Noise in Machine Translation
by: Meng, Yan, et al.
Published: (2024)
by: Meng, Yan, et al.
Published: (2024)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
by: Vashistha, Sachin, et al.
Published: (2025)
by: Vashistha, Sachin, et al.
Published: (2025)
Binary Neural Networks for Large Language Model: A Survey
by: Liu, Liangdong, et al.
Published: (2025)
by: Liu, Liangdong, et al.
Published: (2025)
Learning Disentangled Semantic Spaces of Explanations via Invertible Neural Networks
by: Zhang, Yingji, et al.
Published: (2023)
by: Zhang, Yingji, et al.
Published: (2023)
Persian Pronoun Resolution: Leveraging Neural Networks and Language Models
by: Mohammadi, Hassan Haji, et al.
Published: (2024)
by: Mohammadi, Hassan Haji, et al.
Published: (2024)
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?
by: Lee, Chia-Hsuan, et al.
Published: (2026)
by: Lee, Chia-Hsuan, et al.
Published: (2026)
Measuring Meaning Composition in the Human Brain with Composition Scores from Large Language Models
by: Gao, Changjiang, et al.
Published: (2024)
by: Gao, Changjiang, et al.
Published: (2024)
Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
by: Scivetti, Wesley, et al.
Published: (2025)
by: Scivetti, Wesley, et al.
Published: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Similar Items
-
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
by: Li, Kenneth, et al.
Published: (2024) -
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025) -
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
by: Li, Kenneth, et al.
Published: (2022) -
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
by: Li, Kenneth, et al.
Published: (2023) -
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
by: Li, Kenneth, et al.
Published: (2024)