Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Amos, Ido, Berant, Jonathan, Gupta, Ankit |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
Ensemble Self-Training for Unsupervised Machine Translation
by: Aharon, Ido, et al.
Published: (2026)
by: Aharon, Ido, et al.
Published: (2026)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
by: Eisenstein, Jacob, et al.
Published: (2026)
by: Eisenstein, Jacob, et al.
Published: (2026)
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
by: Wen, Liang, et al.
Published: (2025)
by: Wen, Liang, et al.
Published: (2025)
Sequence-Level Leakage Risk of Training Data in Large Language Models
by: Tiwari, Trishita, et al.
Published: (2024)
by: Tiwari, Trishita, et al.
Published: (2024)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrieval
by: Rubin, Ohad, et al.
Published: (2023)
by: Rubin, Ohad, et al.
Published: (2023)
Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels
by: Xian, Jasper, et al.
Published: (2024)
by: Xian, Jasper, et al.
Published: (2024)
Don't lie to your friends: Learning what you know from collaborative self-play
by: Eisenstein, Jacob, et al.
Published: (2025)
by: Eisenstein, Jacob, et al.
Published: (2025)
SEMQA: Semi-Extractive Multi-Source Question Answering
by: Schuster, Tal, et al.
Published: (2023)
by: Schuster, Tal, et al.
Published: (2023)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025)
by: Mao, Hanyi, et al.
Published: (2025)
QuRating: Selecting High-Quality Data for Training Language Models
by: Wettig, Alexander, et al.
Published: (2024)
by: Wettig, Alexander, et al.
Published: (2024)
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
by: Zou, Haosheng, et al.
Published: (2025)
by: Zou, Haosheng, et al.
Published: (2025)
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
by: Wunderle, Julia, et al.
Published: (2025)
by: Wunderle, Julia, et al.
Published: (2025)
Long-range Modeling and Processing of Multimodal Event Sequences
by: Li, Jichu, et al.
Published: (2026)
by: Li, Jichu, et al.
Published: (2026)
ALTA: Compiler-Based Analysis of Transformers
by: Shaw, Peter, et al.
Published: (2024)
by: Shaw, Peter, et al.
Published: (2024)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
Cascade-Aware Training of Language Models
by: Wang, Congchao, et al.
Published: (2024)
by: Wang, Congchao, et al.
Published: (2024)
LLMs Plagiarize: Ensuring Responsible Sourcing of Large Language Model Training Data Through Knowledge Graph Comparison
by: Mondal, Devam, et al.
Published: (2024)
by: Mondal, Devam, et al.
Published: (2024)
Extending Input Contexts of Language Models through Training on Segmented Sequences
by: Karypis, Petros, et al.
Published: (2023)
by: Karypis, Petros, et al.
Published: (2023)
Training Language Models on Synthetic Edit Sequences Improves Code Synthesis
by: Piterbarg, Ulyana, et al.
Published: (2024)
by: Piterbarg, Ulyana, et al.
Published: (2024)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
by: Setlur, Amrith, et al.
Published: (2024)
by: Setlur, Amrith, et al.
Published: (2024)
How Well Can a Long Sequence Model Model Long Sequences? Comparing Architechtural Inductive Biases on Long-Context Abilities
by: Huang, Jerry
Published: (2024)
by: Huang, Jerry
Published: (2024)
TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning
by: Zheng, Mingyu, et al.
Published: (2025)
by: Zheng, Mingyu, et al.
Published: (2025)
How to Train Long-Context Language Models (Effectively)
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
Anterior's Approach to Fairness Evaluation of Automated Prior Authorization System
by: Selvaraj, Sai P., et al.
Published: (2026)
by: Selvaraj, Sai P., et al.
Published: (2026)
LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch
by: Pfister, Jan, et al.
Published: (2024)
by: Pfister, Jan, et al.
Published: (2024)
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
by: Goren, Shani, et al.
Published: (2026)
by: Goren, Shani, et al.
Published: (2026)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
On Training Data Influence of GPT Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Unveiling the Mystery of Weight in Large Foundation Models: Gaussian Distribution Never Fades
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
by: Yan, Shaotian, et al.
Published: (2026)
by: Yan, Shaotian, et al.
Published: (2026)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
by: Rawat, Ankit Singh, et al.
Published: (2024)
by: Rawat, Ankit Singh, et al.
Published: (2024)
Theoretical guarantees on the best-of-n alignment policy
by: Beirami, Ahmad, et al.
Published: (2024)
by: Beirami, Ahmad, et al.
Published: (2024)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
by: Basu, Soumya, et al.
Published: (2024)
by: Basu, Soumya, et al.
Published: (2024)
Ayn: A Tiny yet Competitive Indian Legal Language Model Pretrained from Scratch
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
Equitable Electronic Health Record Prediction with FAME: Fairness-Aware Multimodal Embedding
by: Hooman, Nikkie, et al.
Published: (2025)
by: Hooman, Nikkie, et al.
Published: (2025)
Cold-Start Personalization via Training-Free Priors from Structured World Models
by: Bose, Avinandan, et al.
Published: (2026)
by: Bose, Avinandan, et al.
Published: (2026)
Similar Items
-
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
by: Wang, Haoming, et al.
Published: (2025) -
Ensemble Self-Training for Unsupervised Machine Translation
by: Aharon, Ido, et al.
Published: (2026) -
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
by: Eisenstein, Jacob, et al.
Published: (2026) -
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
by: Bordt, Sebastian, et al.
Published: (2024) -
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
by: Wen, Liang, et al.
Published: (2025)