Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
Fuente:
arXiv
Saved in:
| Main Authors: | Roger, Alexis, Legate, Gwen, Rasul, Kashif, Nevmyvaka, Yuriy, Rish, Irina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
by: Riachi, Roland, et al.
Published: (2025)
by: Riachi, Roland, et al.
Published: (2025)
LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series
by: Roger, Alexis, et al.
Published: (2026)
by: Roger, Alexis, et al.
Published: (2026)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
by: Garg, Sahil, et al.
Published: (2024)
by: Garg, Sahil, et al.
Published: (2024)
TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
by: Ning, Kanghui, et al.
Published: (2025)
by: Ning, Kanghui, et al.
Published: (2025)
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
by: Rasul, Kashif, et al.
Published: (2023)
by: Rasul, Kashif, et al.
Published: (2023)
Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
by: Vaidhya, Tejas, et al.
Published: (2025)
by: Vaidhya, Tejas, et al.
Published: (2025)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
by: Hogan, Brendan R., et al.
Published: (2026)
by: Hogan, Brendan R., et al.
Published: (2026)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
by: Chen, Daiwei, et al.
Published: (2026)
by: Chen, Daiwei, et al.
Published: (2026)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
by: Ibrahim, Adam, et al.
Published: (2024)
by: Ibrahim, Adam, et al.
Published: (2024)
Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
by: Kaushal, Ayush, et al.
Published: (2024)
by: Kaushal, Ayush, et al.
Published: (2024)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026)
by: Yu, Zichun, et al.
Published: (2026)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
by: Goel, Vatsal, et al.
Published: (2026)
by: Goel, Vatsal, et al.
Published: (2026)
Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
by: Lewis, Ashley, et al.
Published: (2025)
by: Lewis, Ashley, et al.
Published: (2025)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
by: Huang, Ruizhe, et al.
Published: (2025)
by: Huang, Ruizhe, et al.
Published: (2025)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Fast Vocabulary Transfer for Language Model Compression
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Structural Knowledge Informed Continual Multivariate Time Series Forecasting
by: Pan, Zijie, et al.
Published: (2024)
by: Pan, Zijie, et al.
Published: (2024)
Large Language Models for Time Series: A Survey
by: Zhang, Xiyuan, et al.
Published: (2024)
by: Zhang, Xiyuan, et al.
Published: (2024)
In-Context Fine-Tuning for Time-Series Foundation Models
by: Das, Abhimanyu, et al.
Published: (2024)
by: Das, Abhimanyu, et al.
Published: (2024)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025)
by: Huang, Jenny Y., et al.
Published: (2025)
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
by: Hu, Zhimin, et al.
Published: (2026)
by: Hu, Zhimin, et al.
Published: (2026)
In-context Time Series Predictor
by: Lu, Jiecheng, et al.
Published: (2024)
by: Lu, Jiecheng, et al.
Published: (2024)
In-context Pretraining: Language Modeling Beyond Document Boundaries
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
Fine-Tuning a Time Series Foundation Model with Wasserstein Loss
by: Chernov, Andrei
Published: (2024)
by: Chernov, Andrei
Published: (2024)
Text2TimeSeries: Enhancing Financial Forecasting through Time Series Prediction Updates with Event-Driven Insights from Large Language Models
by: Kurisinkel, Litton Jose, et al.
Published: (2024)
by: Kurisinkel, Litton Jose, et al.
Published: (2024)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
by: Bayazit, Deniz, et al.
Published: (2025)
by: Bayazit, Deniz, et al.
Published: (2025)
Too Big to Fool: Resisting Deception in Language Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
TableTime: Reformulating Time Series Classification as Training-Free Table Understanding with Large Language Models
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
Similar Items
-
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
by: Riachi, Roland, et al.
Published: (2025) -
LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series
by: Roger, Alexis, et al.
Published: (2026) -
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025) -
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
by: Garg, Sahil, et al.
Published: (2024) -
TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
by: Ning, Kanghui, et al.
Published: (2025)