Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Mengnan, Huang, Yufan, Yao, Yongqiang, Wang, Maoquan, Gu, Bin, Sundaresan, Neel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
by: Dai, Hankun, et al.
Published: (2025)
by: Dai, Hankun, et al.
Published: (2025)
When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming
by: Corielli, Francesco
Published: (2026)
by: Corielli, Francesco
Published: (2026)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
DeepCircuitX: A Comprehensive Repository-Level Dataset for RTL Code Understanding, Generation, and PPA Analysis
by: Li, Zeju, et al.
Published: (2025)
by: Li, Zeju, et al.
Published: (2025)
Self-Infilling Code Generation
by: Zheng, Lin, et al.
Published: (2023)
by: Zheng, Lin, et al.
Published: (2023)
ENTP: Encoder-only Next Token Prediction
by: Ewer, Ethan, et al.
Published: (2024)
by: Ewer, Ethan, et al.
Published: (2024)
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
RefineStat: Efficient Exploration for Probabilistic Program Synthesis
by: Kanda, Madhav, et al.
Published: (2025)
by: Kanda, Madhav, et al.
Published: (2025)
Type-Constrained Code Generation with Language Models
by: Mündler, Niels, et al.
Published: (2025)
by: Mündler, Niels, et al.
Published: (2025)
Adaptive Branch-and-Bound Tree Exploration for Neural Network Verification
by: Fukuda, Kota, et al.
Published: (2025)
by: Fukuda, Kota, et al.
Published: (2025)
The Next 700 ML-Enabled Compiler Optimizations
by: VenkataKeerthy, S., et al.
Published: (2023)
by: VenkataKeerthy, S., et al.
Published: (2023)
More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
by: Zi, Yangtian, et al.
Published: (2025)
by: Zi, Yangtian, et al.
Published: (2025)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
by: Rohekar, Raanan Y., et al.
Published: (2024)
by: Rohekar, Raanan Y., et al.
Published: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
CompilerDream: Learning a Compiler World Model for General Code Optimization
by: Deng, Chaoyi, et al.
Published: (2024)
by: Deng, Chaoyi, et al.
Published: (2024)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
by: Noah, Amit Finkman, et al.
Published: (2024)
by: Noah, Amit Finkman, et al.
Published: (2024)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
by: Schneider, Johannes
Published: (2024)
by: Schneider, Johannes
Published: (2024)
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
by: Huang, Guanhua, et al.
Published: (2025)
by: Huang, Guanhua, et al.
Published: (2025)
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
AnCoder: Anchored Code Generation via Discrete Diffusion Models
by: Xue, Anton, et al.
Published: (2026)
by: Xue, Anton, et al.
Published: (2026)
Combining Neural Architecture Search and Automatic Code Optimization: A Survey
by: Bachiri, Inas, et al.
Published: (2024)
by: Bachiri, Inas, et al.
Published: (2024)
Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning
by: Qin, Zhanyue, et al.
Published: (2026)
by: Qin, Zhanyue, et al.
Published: (2026)
A Comprehensive Guide to Combining R and Python code for Data Science, Machine Learning and Reinforcement Learning
by: Navarro, Alejandro L. García, et al.
Published: (2024)
by: Navarro, Alejandro L. García, et al.
Published: (2024)
Cost-Driven Synthesis of Sound Abstract Interpreters
by: Gu, Qiuhan, et al.
Published: (2025)
by: Gu, Qiuhan, et al.
Published: (2025)
Explorations of Self-Repair in Language Models
by: Rushing, Cody, et al.
Published: (2024)
by: Rushing, Cody, et al.
Published: (2024)
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
by: Dong, Shouyang, et al.
Published: (2025)
by: Dong, Shouyang, et al.
Published: (2025)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
by: Xia, Chunqiu Steven, et al.
Published: (2024)
by: Xia, Chunqiu Steven, et al.
Published: (2024)
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Synthetic Programming Elicitation for Text-to-Code in Very Low-Resource Programming and Formal Languages
by: Mora, Federico, et al.
Published: (2024)
by: Mora, Federico, et al.
Published: (2024)
Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs
by: Cassano, Federico, et al.
Published: (2023)
by: Cassano, Federico, et al.
Published: (2023)
Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment
by: Boruch-Gruszecki, Aleksander, et al.
Published: (2025)
by: Boruch-Gruszecki, Aleksander, et al.
Published: (2025)
Neural Models for Source Code Synthesis and Completion
by: Niyogi, Mitodru
Published: (2024)
by: Niyogi, Mitodru
Published: (2024)
Code Simulation Challenges for Large Language Models
by: La Malfa, Emanuele, et al.
Published: (2024)
by: La Malfa, Emanuele, et al.
Published: (2024)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
A Multi-Perspective Architecture for Semantic Code Search
by: Haldar, Rajarshi, et al.
Published: (2020)
by: Haldar, Rajarshi, et al.
Published: (2020)
LILO: Learning Interpretable Libraries by Compressing and Documenting Code
by: Grand, Gabriel, et al.
Published: (2023)
by: Grand, Gabriel, et al.
Published: (2023)
A Deep Learning Model for Predicting Transformation Legality
by: Tiwari, Avani, et al.
Published: (2025)
by: Tiwari, Avani, et al.
Published: (2025)
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Similar Items
-
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
by: Dai, Hankun, et al.
Published: (2025) -
When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming
by: Corielli, Francesco
Published: (2026) -
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024) -
DeepCircuitX: A Comprehensive Repository-Level Dataset for RTL Code Understanding, Generation, and PPA Analysis
by: Li, Zeju, et al.
Published: (2025) -
Self-Infilling Code Generation
by: Zheng, Lin, et al.
Published: (2023)