Not all tokens are needed(NAT): token efficient reinforcement learning
Fuente:
arXiv
Saved in:
| Main Authors: | Sang, Hejian, Xu, Yuanda, Zhou, Zhengze, He, Ran, Wang, Zhipeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
TIP: Token Importance in On-Policy Distillation
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
Next-token pretraining implies in-context learning
by: Riechers, Paul M., et al.
Published: (2025)
by: Riechers, Paul M., et al.
Published: (2025)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
Physics in Next-token Prediction
by: An, Hongjun, et al.
Published: (2024)
by: An, Hongjun, et al.
Published: (2024)
Debunk the Myth of SFT Generalization
by: Lin, Xiaofeng, et al.
Published: (2025)
by: Lin, Xiaofeng, et al.
Published: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Shaping capabilities with token-level data filtering
by: Rathi, Neil, et al.
Published: (2026)
by: Rathi, Neil, et al.
Published: (2026)
Aligning Diffusion Language Models via Unpaired Preference Optimization
by: Jindal, Vaibhav, et al.
Published: (2025)
by: Jindal, Vaibhav, et al.
Published: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
by: Bulatov, Aydar, et al.
Published: (2023)
by: Bulatov, Aydar, et al.
Published: (2023)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
Bio2Token: All-atom tokenization of any biomolecular structure with Mamba
by: Liu, Andrew, et al.
Published: (2024)
by: Liu, Andrew, et al.
Published: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
by: Sokar, Ghada, et al.
Published: (2024)
by: Sokar, Ghada, et al.
Published: (2024)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
by: Kwek, Eugene, et al.
Published: (2025)
by: Kwek, Eugene, et al.
Published: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
by: Marconato, Emanuele, et al.
Published: (2024)
by: Marconato, Emanuele, et al.
Published: (2024)
Essential-Web v1.0: 24T tokens of organized web data
by: AI, Essential, et al.
Published: (2025)
by: AI, Essential, et al.
Published: (2025)
Fast, memory-efficient genomic interval tokenizers for modern machine learning
by: LeRoy, Nathan J., et al.
Published: (2025)
by: LeRoy, Nathan J., et al.
Published: (2025)
LookupViT: Compressing visual information to a limited number of tokens
by: Koner, Rajat, et al.
Published: (2024)
by: Koner, Rajat, et al.
Published: (2024)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
GReF: A Unified Generative Framework for Efficient Reranking via Ordered Multi-token Prediction
by: Lin, Zhijie, et al.
Published: (2025)
by: Lin, Zhijie, et al.
Published: (2025)
An efficient deep reinforcement learning environment for flexible job-shop scheduling
by: Wu, Xinquan, et al.
Published: (2025)
by: Wu, Xinquan, et al.
Published: (2025)
Designing an efficient and equitable humanitarian supply chain dynamically via reinforcement learning
by: Jin, Weijia
Published: (2025)
by: Jin, Weijia
Published: (2025)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
by: Lin, Xiaofeng, et al.
Published: (2026)
by: Lin, Xiaofeng, et al.
Published: (2026)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
Adaptive traffic signal safety and efficiency improvement by multi objective deep reinforcement learning approach
by: Mirbakhsh, Shahin, et al.
Published: (2024)
by: Mirbakhsh, Shahin, et al.
Published: (2024)
Tabular Data: Is Deep Learning all you need?
by: Zabërgja, Guri, et al.
Published: (2024)
by: Zabërgja, Guri, et al.
Published: (2024)
Curriculum reinforcement learning with measurable task representation learning
by: Wen, Yongyan, et al.
Published: (2026)
by: Wen, Yongyan, et al.
Published: (2026)
Scilab-RL: A software framework for efficient reinforcement learning and cognitive modeling research
by: Dohmen, Jan, et al.
Published: (2024)
by: Dohmen, Jan, et al.
Published: (2024)
Discovering highly efficient low-weight quantum error-correcting codes with reinforcement learning
by: He, Austin Yubo, et al.
Published: (2025)
by: He, Austin Yubo, et al.
Published: (2025)
Normalization and effective learning rates in reinforcement learning
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
An approach of deep reinforcement learning for maximizing the net present value of stochastic projects
by: Xu, Wei, et al.
Published: (2025)
by: Xu, Wei, et al.
Published: (2025)
Similar Items
-
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
by: Xu, Yuanda, et al.
Published: (2026) -
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026) -
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026) -
TIP: Token Importance in On-Policy Distillation
by: Xu, Yuanda, et al.
Published: (2026) -
CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
by: Sang, Hejian, et al.
Published: (2026)