Efficient Training of Language Models with Compact and Consistent Next Token Distributions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sathe, Ashutosh, Sarawagi, Sunita |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Self Consistency in LLMs through Probabilistic Tokenization
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
Robust Root Cause Diagnosis using In-Distribution Interventions
von: Nagalapatti, Lokesh, et al.
Veröffentlicht: (2025)
von: Nagalapatti, Lokesh, et al.
Veröffentlicht: (2025)
SALSA: Speedy ASR-LLM Synchronous Aggregation
von: Mittal, Ashish, et al.
Veröffentlicht: (2024)
von: Mittal, Ashish, et al.
Veröffentlicht: (2024)
Reasoning Bias of Next Token Prediction Training
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025)
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025)
Temporally Consistent Factuality Probing for Large Language Models
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
TFMAdapter: Lightweight Instance-Level Adaptation of Foundation Models for Forecasting with Covariates
von: Dange, Afrin, et al.
Veröffentlicht: (2025)
von: Dange, Afrin, et al.
Veröffentlicht: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
Multimodal Latent Language Modeling with Next-Token Diffusion
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
von: Wu, Chen, et al.
Veröffentlicht: (2025)
von: Wu, Chen, et al.
Veröffentlicht: (2025)
A Law of Next-Token Prediction in Large Language Models
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
Differentially Private Next-Token Prediction of Large Language Models
von: Flemings, James, et al.
Veröffentlicht: (2024)
von: Flemings, James, et al.
Veröffentlicht: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
Masked Diffusion Models are Secretly Learned-Order Autoregressive Models
von: Garg, Prateek, et al.
Veröffentlicht: (2025)
von: Garg, Prateek, et al.
Veröffentlicht: (2025)
DiJiang: Efficient Large Language Models through Compact Kernelization
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
Efficient Temporal Tokenization for Mobility Prediction with Large Language Models
von: He, Haoyu, et al.
Veröffentlicht: (2025)
von: He, Haoyu, et al.
Veröffentlicht: (2025)
From Search To Sampling: Generative Models For Robust Algorithmic Recourse
von: Garg, Prateek, et al.
Veröffentlicht: (2025)
von: Garg, Prateek, et al.
Veröffentlicht: (2025)
ENTP: Encoder-only Next Token Prediction
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
Training Language Models to Reason Efficiently
von: Arora, Daman, et al.
Veröffentlicht: (2025)
von: Arora, Daman, et al.
Veröffentlicht: (2025)
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
von: Tran, Toan, et al.
Veröffentlicht: (2025)
von: Tran, Toan, et al.
Veröffentlicht: (2025)
Training In-Context and In-Weights Mixtures Via Contrastive Context Sampling
von: Malu, Deeptanshu, et al.
Veröffentlicht: (2026)
von: Malu, Deeptanshu, et al.
Veröffentlicht: (2026)
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023)
von: Malach, Eran
Veröffentlicht: (2023)
Cross-Language Assessment of Mathematical Capability of ChatGPT
von: Sathe, Gargi, et al.
Veröffentlicht: (2024)
von: Sathe, Gargi, et al.
Veröffentlicht: (2024)
PairNet: Training with Observed Pairs to Estimate Individual Treatment Effect
von: Nagalapatti, Lokesh, et al.
Veröffentlicht: (2024)
von: Nagalapatti, Lokesh, et al.
Veröffentlicht: (2024)
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space
von: Rom, Aviad, et al.
Veröffentlicht: (2024)
von: Rom, Aviad, et al.
Veröffentlicht: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
von: Plyler, Mitchell, et al.
Veröffentlicht: (2025)
von: Plyler, Mitchell, et al.
Veröffentlicht: (2025)
Better Language Model Inversion by Compactly Representing Next-Token Distributions
von: Nazir, Murtaza, et al.
Veröffentlicht: (2025)
von: Nazir, Murtaza, et al.
Veröffentlicht: (2025)
Token-Efficient Leverage Learning in Large Language Models
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2024)
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2024)
Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
von: Qi, Mengnan, et al.
Veröffentlicht: (2024)
von: Qi, Mengnan, et al.
Veröffentlicht: (2024)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
von: Schneider, Johannes
Veröffentlicht: (2024)
von: Schneider, Johannes
Veröffentlicht: (2024)
Language Modeling with Learned Meta-Tokens
von: Shah, Alok N., et al.
Veröffentlicht: (2025)
von: Shah, Alok N., et al.
Veröffentlicht: (2025)
Parallel Token Prediction for Language Models
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Improving Self Consistency in LLMs through Probabilistic Tokenization
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024) -
Robust Root Cause Diagnosis using In-Distribution Interventions
von: Nagalapatti, Lokesh, et al.
Veröffentlicht: (2025) -
SALSA: Speedy ASR-LLM Synchronous Aggregation
von: Mittal, Ashish, et al.
Veröffentlicht: (2024) -
Reasoning Bias of Next Token Prediction Training
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025) -
Temporally Consistent Factuality Probing for Large Language Models
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024)