Efficient Joint Prediction of Multiple Future Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ahn, Kwangjun, Lamb, Alex, Langford, John |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Belief State Transformer
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
TokenButler: Token Importance is Predictable
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
von: Yang, Kaisen, et al.
Veröffentlicht: (2025)
von: Yang, Kaisen, et al.
Veröffentlicht: (2025)
Dion2: A Simple Method to Shrink Matrix in Muon
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
Towards Principled Representation Learning from Videos for Reinforcement Learning
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
von: Chatzi, Ivi, et al.
Veröffentlicht: (2025)
von: Chatzi, Ivi, et al.
Veröffentlicht: (2025)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
von: Shin, Seungjun, et al.
Veröffentlicht: (2025)
von: Shin, Seungjun, et al.
Veröffentlicht: (2025)
Lossless Token Sequence Compression via Meta-Tokens
von: Harvill, John, et al.
Veröffentlicht: (2025)
von: Harvill, John, et al.
Veröffentlicht: (2025)
Token-Efficient Leverage Learning in Large Language Models
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2024)
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2024)
Multiplicative-Additive Constrained Models:Toward Joint Visualization of Interactive and Independent Effects
von: Wang, Fumin
Veröffentlicht: (2025)
von: Wang, Fumin
Veröffentlicht: (2025)
A Law of Next-Token Prediction in Large Language Models
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
von: Jia, Mumin, et al.
Veröffentlicht: (2025)
von: Jia, Mumin, et al.
Veröffentlicht: (2025)
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models
von: Kongmanee, Jaturong
Veröffentlicht: (2025)
von: Kongmanee, Jaturong
Veröffentlicht: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
von: Cao, Qi, et al.
Veröffentlicht: (2026)
von: Cao, Qi, et al.
Veröffentlicht: (2026)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
von: Zhong, Qimin, et al.
Veröffentlicht: (2026)
von: Zhong, Qimin, et al.
Veröffentlicht: (2026)
Federated Learning-Enabled Hybrid Language Models for Communication-Efficient Token Transmission
von: Solat, Faranaksadat, et al.
Veröffentlicht: (2025)
von: Solat, Faranaksadat, et al.
Veröffentlicht: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
The Foundations of Tokenization: Statistical and Computational Concerns
von: Gastaldi, Juan Luis, et al.
Veröffentlicht: (2024)
von: Gastaldi, Juan Luis, et al.
Veröffentlicht: (2024)
Learning to Route LLMs with Confidence Tokens
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
Dion: Distributed Orthonormalized Updates
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
Adversarial Tokenization
von: Geh, Renato Lui, et al.
Veröffentlicht: (2025)
von: Geh, Renato Lui, et al.
Veröffentlicht: (2025)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
Understanding World or Predicting Future? A Comprehensive Survey of World Models
von: Ding, Jingtao, et al.
Veröffentlicht: (2024)
von: Ding, Jingtao, et al.
Veröffentlicht: (2024)
How to escape sharp minima with random perturbations
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Belief State Transformer
von: Hu, Edward S., et al.
Veröffentlicht: (2024) -
TokenButler: Token Importance is Predictable
von: Akhauri, Yash, et al.
Veröffentlicht: (2025) -
Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
von: Yang, Kaisen, et al.
Veröffentlicht: (2025) -
Dion2: A Simple Method to Shrink Matrix in Muon
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025) -
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)