Saved in:
| Main Authors: | Tian, Zhuojing, Chen, Yushu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.21448 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fairness under uncertainty in sequential decisions
by: Lee, Michelle Seng Ah, et al.
Published: (2026)
by: Lee, Michelle Seng Ah, et al.
Published: (2026)
Optimal sequential decision-making for error propagation mitigation in digital twins
by: Najafi, Annice, et al.
Published: (2026)
by: Najafi, Annice, et al.
Published: (2026)
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty
by: Martins, Jonas Mayer, et al.
Published: (2026)
by: Martins, Jonas Mayer, et al.
Published: (2026)
Mathematics of statistical sequential decision-making: concentration, risk-awareness and modelling in stochastic bandits, with applications to bariatric surgery
by: Saux, Patrick
Published: (2024)
by: Saux, Patrick
Published: (2024)
Counterfactual inference in sequential experiments
by: Dwivedi, Raaz, et al.
Published: (2022)
by: Dwivedi, Raaz, et al.
Published: (2022)
Visualizing token importance for black-box language models
by: Rauba, Paulius, et al.
Published: (2025)
by: Rauba, Paulius, et al.
Published: (2025)
Do language models plan ahead for future tokens?
by: Wu, Wilson, et al.
Published: (2024)
by: Wu, Wilson, et al.
Published: (2024)
Quasi-Bayes empirical Bayes: a sequential approach to the Poisson compound decision problem
by: Favaro, Stefano, et al.
Published: (2024)
by: Favaro, Stefano, et al.
Published: (2024)
Efficient numeracy in language models through single-token number embeddings
by: Kreitner, Linus, et al.
Published: (2025)
by: Kreitner, Linus, et al.
Published: (2025)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
Physics in Next-token Prediction
by: An, Hongjun, et al.
Published: (2024)
by: An, Hongjun, et al.
Published: (2024)
On convex decision regions in deep network representations
by: Tětková, Lenka, et al.
Published: (2023)
by: Tětková, Lenka, et al.
Published: (2023)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
EvoMerge: Neuroevolution for Large Language Models
by: Jiang, Yushu
Published: (2024)
by: Jiang, Yushu
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
The cell as a token: high-dimensional geometry in language models and cell embeddings
by: Gilpin, William
Published: (2025)
by: Gilpin, William
Published: (2025)
Tokenphormer: Structure-aware Multi-token Graph Transformer for Node Classification
by: Zhou, Zijie, et al.
Published: (2024)
by: Zhou, Zijie, et al.
Published: (2024)
GReF: A Unified Generative Framework for Efficient Reranking via Ordered Multi-token Prediction
by: Lin, Zhijie, et al.
Published: (2025)
by: Lin, Zhijie, et al.
Published: (2025)
Long-range gene expression prediction with token alignment of large language model
by: Honig, Edouardo, et al.
Published: (2024)
by: Honig, Edouardo, et al.
Published: (2024)
Next-token pretraining implies in-context learning
by: Riechers, Paul M., et al.
Published: (2025)
by: Riechers, Paul M., et al.
Published: (2025)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
by: Huang, Ruiquan, et al.
Published: (2024)
by: Huang, Ruiquan, et al.
Published: (2024)
Counterfactual Image Editing
by: Pan, Yushu, et al.
Published: (2024)
by: Pan, Yushu, et al.
Published: (2024)
Comparison of Autoencoders for tokenization of ASL datasets
by: Praun-Petrovic, Vouk, et al.
Published: (2025)
by: Praun-Petrovic, Vouk, et al.
Published: (2025)
Bayesian full waveform inversion with sequential surrogate model refinement
by: Meles, Giovanni Angelo, et al.
Published: (2025)
by: Meles, Giovanni Angelo, et al.
Published: (2025)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Increasing transformer token length with a Maximum Entropy Principle Method
by: Cukier, R. I.
Published: (2024)
by: Cukier, R. I.
Published: (2024)
A two-step sequential approach for hyperparameter selection in finite context models
by: Contente, José, et al.
Published: (2026)
by: Contente, José, et al.
Published: (2026)
From Black-box to Causal-box: Towards Building More Interpretable Models
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
Shaping capabilities with token-level data filtering
by: Rathi, Neil, et al.
Published: (2026)
by: Rathi, Neil, et al.
Published: (2026)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
Quasi-Bayesian sequential deconvolution
by: Favaro, Stefano, et al.
Published: (2024)
by: Favaro, Stefano, et al.
Published: (2024)
Fast, memory-efficient genomic interval tokenizers for modern machine learning
by: LeRoy, Nathan J., et al.
Published: (2025)
by: LeRoy, Nathan J., et al.
Published: (2025)
Interpretable epistemic uncertainty decomposition in sequential generative models via polynomial chaos surrogates
by: Nartallo-Kaluarachchi, Ramón, et al.
Published: (2025)
by: Nartallo-Kaluarachchi, Ramón, et al.
Published: (2025)
Gearing Gaussian process modeling and sequential design towards stochastic simulators
by: Binois, Mickael, et al.
Published: (2024)
by: Binois, Mickael, et al.
Published: (2024)
Scaling Transformer to 1M tokens and beyond with RMT
by: Bulatov, Aydar, et al.
Published: (2023)
by: Bulatov, Aydar, et al.
Published: (2023)
Similar Items
-
Fairness under uncertainty in sequential decisions
by: Lee, Michelle Seng Ah, et al.
Published: (2026) -
Optimal sequential decision-making for error propagation mitigation in digital twins
by: Najafi, Annice, et al.
Published: (2026) -
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026) -
What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty
by: Martins, Jonas Mayer, et al.
Published: (2026) -
Mathematics of statistical sequential decision-making: concentration, risk-awareness and modelling in stochastic bandits, with applications to bariatric surgery
by: Saux, Patrick
Published: (2024)