A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Rohekar, Raanan Y., Gurwicz, Yaniv, Yu, Sungduk, Aflalo, Estelle, Lal, Vasudev |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
par: Yu, Sungduk, et autres
Publié: (2025)
par: Yu, Sungduk, et autres
Publié: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
par: Yu, Sungduk, et autres
Publié: (2024)
par: Yu, Sungduk, et autres
Publié: (2024)
ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution
par: Yu, Sungduk, et autres
Publié: (2024)
par: Yu, Sungduk, et autres
Publié: (2024)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024)
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024)
Learning from Reasoning Failures via Synthetic Data Generation
par: Stan, Gabriela Ben Melech, et autres
Publié: (2025)
par: Stan, Gabriela Ben Melech, et autres
Publié: (2025)
Probing Semantic Routing in Large Mixture-of-Expert Models
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
par: Madasu, Avinash, et autres
Publié: (2023)
par: Madasu, Avinash, et autres
Publié: (2023)
Exploring Next Token Prediction in Theory of Mind (ToM) Tasks: Comparative Experiments with GPT-2 and LLaMA-2 AI Models
par: Yadav, Pavan, et autres
Publié: (2025)
par: Yadav, Pavan, et autres
Publié: (2025)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
par: Kundu, Souvik, et autres
Publié: (2025)
par: Kundu, Souvik, et autres
Publié: (2025)
Cautious Next Token Prediction
par: Wang, Yizhou, et autres
Publié: (2025)
par: Wang, Yizhou, et autres
Publié: (2025)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
par: Madasu, Avinash, et autres
Publié: (2025)
par: Madasu, Avinash, et autres
Publié: (2025)
FastRM: An efficient and automatic explainability framework for multimodal generative models
par: Stan, Gabriela Ben-Melech, et autres
Publié: (2024)
par: Stan, Gabriela Ben-Melech, et autres
Publié: (2024)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
par: Hinck, Musashi, et autres
Publié: (2024)
par: Hinck, Musashi, et autres
Publié: (2024)
Steering Large Language Models to Evaluate and Amplify Creativity
par: Olson, Matthew Lyle, et autres
Publié: (2024)
par: Olson, Matthew Lyle, et autres
Publié: (2024)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
par: Alabdulmohsin, Ibrahim, et autres
Publié: (2024)
par: Alabdulmohsin, Ibrahim, et autres
Publié: (2024)
Alternatives To Next Token Prediction In Text Generation -- A Survey
par: Wyatt, Charlie, et autres
Publié: (2025)
par: Wyatt, Charlie, et autres
Publié: (2025)
A Law of Next-Token Prediction in Large Language Models
par: He, Hangfeng, et autres
Publié: (2024)
par: He, Hangfeng, et autres
Publié: (2024)
Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
par: Yang, Chun-Hao, et autres
Publié: (2025)
par: Yang, Chun-Hao, et autres
Publié: (2025)
Modeling Next-Token Prediction as Left-Nested Intuitionistic Implication
par: Tarau, Paul
Publié: (2026)
par: Tarau, Paul
Publié: (2026)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
par: Aflalo, Estelle, et autres
Publié: (2024)
par: Aflalo, Estelle, et autres
Publié: (2024)
Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators
par: Rhee, Phill Kyu
Publié: (2025)
par: Rhee, Phill Kyu
Publié: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
par: Shao, Chenze, et autres
Publié: (2024)
par: Shao, Chenze, et autres
Publié: (2024)
Advancing Pancreatic Cancer Prediction with a Next Visit Token Prediction Head on top of Med-BERT
par: He, Jianping, et autres
Publié: (2025)
par: He, Jianping, et autres
Publié: (2025)
LLMs are Not Just Next Token Predictors
par: Downes, Stephen M., et autres
Publié: (2024)
par: Downes, Stephen M., et autres
Publié: (2024)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
par: Gohil, Vasudev
Publié: (2025)
par: Gohil, Vasudev
Publié: (2025)
ADAM: An Embodied Causal Agent in Open-World Environments
par: Yu, Shu, et autres
Publié: (2024)
par: Yu, Shu, et autres
Publié: (2024)
Mechanics of Next Token Prediction with Self-Attention
par: Li, Yingcong, et autres
Publié: (2024)
par: Li, Yingcong, et autres
Publié: (2024)
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
par: An, Chenyang, et autres
Publié: (2024)
par: An, Chenyang, et autres
Publié: (2024)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
par: Olson, Matthew Lyle, et autres
Publié: (2026)
par: Olson, Matthew Lyle, et autres
Publié: (2026)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
par: Jia, Mumin, et autres
Publié: (2025)
par: Jia, Mumin, et autres
Publié: (2025)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
par: Kim, Minseo, et autres
Publié: (2025)
par: Kim, Minseo, et autres
Publié: (2025)
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
par: Cheng, Yu-Ang, et autres
Publié: (2025)
par: Cheng, Yu-Ang, et autres
Publié: (2025)
From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
par: Tian, Yuchuan, et autres
Publié: (2025)
par: Tian, Yuchuan, et autres
Publié: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
par: Mishra, Shubhra, et autres
Publié: (2024)
par: Mishra, Shubhra, et autres
Publié: (2024)
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges
par: Huang, Yizheng, et autres
Publié: (2024)
par: Huang, Yizheng, et autres
Publié: (2024)
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
par: Zhuang, Haomin, et autres
Publié: (2025)
par: Zhuang, Haomin, et autres
Publié: (2025)
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
par: Zhong, Qimin, et autres
Publié: (2026)
par: Zhong, Qimin, et autres
Publié: (2026)
TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
par: Lu, Pengqian, et autres
Publié: (2025)
par: Lu, Pengqian, et autres
Publié: (2025)
SpecFuse: Ensembling Large Language Models via Next-Segment Prediction
par: Lv, Bo, et autres
Publié: (2024)
par: Lv, Bo, et autres
Publié: (2024)
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
par: Lee, Isabelle, et autres
Publié: (2024)
par: Lee, Isabelle, et autres
Publié: (2024)
Documents similaires
-
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
par: Yu, Sungduk, et autres
Publié: (2025) -
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
par: Yu, Sungduk, et autres
Publié: (2024) -
ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution
par: Yu, Sungduk, et autres
Publié: (2024) -
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024) -
Learning from Reasoning Failures via Synthetic Data Generation
par: Stan, Gabriela Ben Melech, et autres
Publié: (2025)