How Transformers Learn to Plan via Multi-Token Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Jianhao, Zhou, Zhanpeng, Xia, Renqiu, Mirzasoleiman, Baharan, Su, Weijie, Huang, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
di: Huang, Jianhao, et al.
Pubblicazione: (2026)
di: Huang, Jianhao, et al.
Pubblicazione: (2026)
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
di: Joshi, Siddharth, et al.
Pubblicazione: (2023)
di: Joshi, Siddharth, et al.
Pubblicazione: (2023)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
di: Xue, Yihao, et al.
Pubblicazione: (2025)
di: Xue, Yihao, et al.
Pubblicazione: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
di: Javanmard, Adel, et al.
Pubblicazione: (2026)
di: Javanmard, Adel, et al.
Pubblicazione: (2026)
Understanding the Role of Training Data in Test-Time Scaling
di: Javanmard, Adel, et al.
Pubblicazione: (2025)
di: Javanmard, Adel, et al.
Pubblicazione: (2025)
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
di: Gangavarapu, Tushaar, et al.
Pubblicazione: (2026)
di: Gangavarapu, Tushaar, et al.
Pubblicazione: (2026)
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
di: Yang, Yu, et al.
Pubblicazione: (2024)
di: Yang, Yu, et al.
Pubblicazione: (2024)
Graph Contrastive Learning under Heterophily via Graph Filters
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
di: Huang, Jianhao, et al.
Pubblicazione: (2025)
di: Huang, Jianhao, et al.
Pubblicazione: (2025)
Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution Generalization
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
di: Zhong, Qimin, et al.
Pubblicazione: (2025)
di: Zhong, Qimin, et al.
Pubblicazione: (2025)
Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision
di: Xue, Yihao, et al.
Pubblicazione: (2026)
di: Xue, Yihao, et al.
Pubblicazione: (2026)
A Law of Next-Token Prediction in Large Language Models
di: He, Hangfeng, et al.
Pubblicazione: (2024)
di: He, Hangfeng, et al.
Pubblicazione: (2024)
LoRA is All You Need for Safety Alignment of Reasoning LLMs
di: Xue, Yihao, et al.
Pubblicazione: (2025)
di: Xue, Yihao, et al.
Pubblicazione: (2025)
Efficient Real-Time Aircraft ETA Prediction via Feature Tokenization Transformer
di: Huang, Liping, et al.
Pubblicazione: (2025)
di: Huang, Liping, et al.
Pubblicazione: (2025)
Length-MAX Tokenizer for Language Models
di: Dong, Dong, et al.
Pubblicazione: (2025)
di: Dong, Dong, et al.
Pubblicazione: (2025)
LoLaFL: Low-Latency Federated Learning via Forward-only Propagation
di: Zhang, Jierui, et al.
Pubblicazione: (2024)
di: Zhang, Jierui, et al.
Pubblicazione: (2024)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-Training of Deep Networks
di: Joshi, Siddharth, et al.
Pubblicazione: (2024)
di: Joshi, Siddharth, et al.
Pubblicazione: (2024)
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
di: Xue, Yihao, et al.
Pubblicazione: (2023)
di: Xue, Yihao, et al.
Pubblicazione: (2023)
Planning Transformer: Long-Horizon Offline Reinforcement Learning with Planning Tokens
di: Clinton, Joseph, et al.
Pubblicazione: (2024)
di: Clinton, Joseph, et al.
Pubblicazione: (2024)
Parking Availability Prediction via Fusing Multi-Source Data with A Self-Supervised Learning Enhanced Spatio-Temporal Inverted Transformer
di: Huang, Yin, et al.
Pubblicazione: (2025)
di: Huang, Yin, et al.
Pubblicazione: (2025)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
di: Su, Weijie
Pubblicazione: (2025)
di: Su, Weijie
Pubblicazione: (2025)
Adaptive Computation Depth via Learned Token Routing in Transformers
di: Mohammed, Ahmed Abdelmuniem Abdalla
Pubblicazione: (2026)
di: Mohammed, Ahmed Abdelmuniem Abdalla
Pubblicazione: (2026)
Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap
di: Yang, Wenhan, et al.
Pubblicazione: (2025)
di: Yang, Wenhan, et al.
Pubblicazione: (2025)
EasyST: A Simple Framework for Spatio-Temporal Prediction
di: Tang, Jiabin, et al.
Pubblicazione: (2024)
di: Tang, Jiabin, et al.
Pubblicazione: (2024)
Investigating the Impact of Model Width and Density on Generalization in Presence of Label Noise
di: Xue, Yihao, et al.
Pubblicazione: (2022)
di: Xue, Yihao, et al.
Pubblicazione: (2022)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
di: Zhou, Jingbo, et al.
Pubblicazione: (2026)
di: Zhou, Jingbo, et al.
Pubblicazione: (2026)
Adaptive Graph Learning with Transformer for Multi-Reservoir Inflow Prediction
di: Hu, Pengfei, et al.
Pubblicazione: (2025)
di: Hu, Pengfei, et al.
Pubblicazione: (2025)
AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime
di: Su, Jianhao, et al.
Pubblicazione: (2026)
di: Su, Jianhao, et al.
Pubblicazione: (2026)
OptionZero: Planning with Learned Options
di: Huang, Po-Wei, et al.
Pubblicazione: (2025)
di: Huang, Po-Wei, et al.
Pubblicazione: (2025)
Learning with Foresight: Enhancing Neural Routing Policy via Multi-Node Lookahead Prediction
di: Jiang, Xia, et al.
Pubblicazione: (2026)
di: Jiang, Xia, et al.
Pubblicazione: (2026)
Multi-Path Collaborative Reasoning via Reinforcement Learning
di: Lv, Jindi, et al.
Pubblicazione: (2025)
di: Lv, Jindi, et al.
Pubblicazione: (2025)
Global-Lens Transformers: Adaptive Token Mixing for Dynamic Link Prediction
di: Zou, Tao, et al.
Pubblicazione: (2025)
di: Zou, Tao, et al.
Pubblicazione: (2025)
TART: Token-based Architecture Transformer for Neural Network Performance Prediction
di: He, Yannis Y.
Pubblicazione: (2025)
di: He, Yannis Y.
Pubblicazione: (2025)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
Scalable Heterogeneous Graph Learning via Heterogeneous-aware Orthogonal Prototype Experts
di: Zhou, Wei, et al.
Pubblicazione: (2026)
di: Zhou, Wei, et al.
Pubblicazione: (2026)
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
di: Zhong, Qimin, et al.
Pubblicazione: (2026)
di: Zhong, Qimin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
di: Huang, Jianhao, et al.
Pubblicazione: (2026) -
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
di: Joshi, Siddharth, et al.
Pubblicazione: (2023) -
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
di: Xue, Yihao, et al.
Pubblicazione: (2025) -
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
di: Javanmard, Adel, et al.
Pubblicazione: (2026) -
Understanding the Role of Training Data in Test-Time Scaling
di: Javanmard, Adel, et al.
Pubblicazione: (2025)