Guardado en:
| Autores principales: | Jelassi, Samy, Mohri, Clara, Brandfonbrener, David, Gu, Alex, Vyas, Nikhil, Anand, Nikhil, Alvarez-Melis, David, Li, Yuanzhi, Kakade, Sham M., Malach, Eran |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2410.19034 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Repeat After Me: Transformers are Better than State Space Models at Copying
por: Jelassi, Samy, et al.
Publicado: (2024)
por: Jelassi, Samy, et al.
Publicado: (2024)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
por: Brandfonbrener, David, et al.
Publicado: (2024)
por: Brandfonbrener, David, et al.
Publicado: (2024)
Universal Length Generalization with Turing Programs
por: Hou, Kaiying, et al.
Publicado: (2024)
por: Hou, Kaiying, et al.
Publicado: (2024)
The Role of Sparsity for Length Generalization in Transformers
por: Golowich, Noah, et al.
Publicado: (2025)
por: Golowich, Noah, et al.
Publicado: (2025)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
por: Prabhakar, Akshara, et al.
Publicado: (2024)
por: Prabhakar, Akshara, et al.
Publicado: (2024)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
por: Qin, Tian, et al.
Publicado: (2025)
por: Qin, Tian, et al.
Publicado: (2025)
Deconstructing What Makes a Good Optimizer for Language Models
por: Zhao, Rosie, et al.
Publicado: (2024)
por: Zhao, Rosie, et al.
Publicado: (2024)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
por: Zhao, Rosie, et al.
Publicado: (2025)
por: Zhao, Rosie, et al.
Publicado: (2025)
A New Perspective on Shampoo's Preconditioner
por: Morwani, Depen, et al.
Publicado: (2024)
por: Morwani, Depen, et al.
Publicado: (2024)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
por: Li, Kenneth, et al.
Publicado: (2024)
por: Li, Kenneth, et al.
Publicado: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
por: Liu, Bingbin, et al.
Publicado: (2025)
por: Liu, Bingbin, et al.
Publicado: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
por: Vyas, Nikhil, et al.
Publicado: (2024)
por: Vyas, Nikhil, et al.
Publicado: (2024)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
por: Morwani, Depen, et al.
Publicado: (2025)
por: Morwani, Depen, et al.
Publicado: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
por: Abreu, Natalie, et al.
Publicado: (2025)
por: Abreu, Natalie, et al.
Publicado: (2025)
GQ-VAE: A gated quantized VAE for learning variable length tokens
por: Datta, Theo, et al.
Publicado: (2025)
por: Datta, Theo, et al.
Publicado: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
por: Mirtaheri, Parsa, et al.
Publicado: (2025)
por: Mirtaheri, Parsa, et al.
Publicado: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
por: Qin, Tian, et al.
Publicado: (2025)
por: Qin, Tian, et al.
Publicado: (2025)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
por: Su, Huangyuan, et al.
Publicado: (2025)
por: Su, Huangyuan, et al.
Publicado: (2025)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
por: Jelassi, Samy, et al.
Publicado: (2026)
por: Jelassi, Samy, et al.
Publicado: (2026)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
por: Vyas, Nikhil, et al.
Publicado: (2023)
por: Vyas, Nikhil, et al.
Publicado: (2023)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
por: Oncescu, Costin-Andrei, et al.
Publicado: (2026)
por: Oncescu, Costin-Andrei, et al.
Publicado: (2026)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
por: Brandfonbrener, David, et al.
Publicado: (2024)
por: Brandfonbrener, David, et al.
Publicado: (2024)
Transcendence: Generative Models Can Outperform The Experts That Train Them
por: Zhang, Edwin, et al.
Publicado: (2024)
por: Zhang, Edwin, et al.
Publicado: (2024)
How Does Overparameterization Affect Features?
por: Duzgun, Ahmet Cagri, et al.
Publicado: (2024)
por: Duzgun, Ahmet Cagri, et al.
Publicado: (2024)
Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
por: Kim, Jaeyeon, et al.
Publicado: (2026)
por: Kim, Jaeyeon, et al.
Publicado: (2026)
Random Scaling of Emergent Capabilities
por: Zhao, Rosie, et al.
Publicado: (2025)
por: Zhao, Rosie, et al.
Publicado: (2025)
LOTION: Smoothing the Optimization Landscape for Quantized Training
por: Kwun, Mujin, et al.
Publicado: (2025)
por: Kwun, Mujin, et al.
Publicado: (2025)
Auto-Regressive Next-Token Predictors are Universal Learners
por: Malach, Eran
Publicado: (2023)
por: Malach, Eran
Publicado: (2023)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
por: Bansal, Rachit, et al.
Publicado: (2025)
por: Bansal, Rachit, et al.
Publicado: (2025)
How Does Critical Batch Size Scale in Pre-training?
por: Zhang, Hanlin, et al.
Publicado: (2024)
por: Zhang, Hanlin, et al.
Publicado: (2024)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
por: Karchmer, Ari, et al.
Publicado: (2025)
por: Karchmer, Ari, et al.
Publicado: (2025)
Budgeted Multiple-Expert Deferral
por: DeSalvo, Giulia, et al.
Publicado: (2025)
por: DeSalvo, Giulia, et al.
Publicado: (2025)
Hydraulic City
por: Anand, Nikhil
Publicado: (2017)
por: Anand, Nikhil
Publicado: (2017)
Collective Model Intelligence Requires Compatible Specialization
por: Pari, Jyothish, et al.
Publicado: (2024)
por: Pari, Jyothish, et al.
Publicado: (2024)
Understanding the Role of Functional Diversity in Weight-Ensembling with Ingredient Selection and Multidimensional Scaling
por: Rojas, Alex, et al.
Publicado: (2024)
por: Rojas, Alex, et al.
Publicado: (2024)
Quasi-Linear Size PCPs with Small Soundness from HDX
por: Bafna, Mitali, et al.
Publicado: (2024)
por: Bafna, Mitali, et al.
Publicado: (2024)
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
por: Bhattacharya, Antara Raaghavi, et al.
Publicado: (2025)
por: Bhattacharya, Antara Raaghavi, et al.
Publicado: (2025)
A Taxonomy of Transcendence
por: Abreu, Natalie, et al.
Publicado: (2025)
por: Abreu, Natalie, et al.
Publicado: (2025)
LLM Priors for ERM over Programs
por: Singhal, Shivam, et al.
Publicado: (2025)
por: Singhal, Shivam, et al.
Publicado: (2025)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
por: Tsilivis, Nikolaos, et al.
Publicado: (2025)
por: Tsilivis, Nikolaos, et al.
Publicado: (2025)
Ejemplares similares
-
Repeat After Me: Transformers are Better than State Space Models at Copying
por: Jelassi, Samy, et al.
Publicado: (2024) -
Loss-to-Loss Prediction: Scaling Laws for All Datasets
por: Brandfonbrener, David, et al.
Publicado: (2024) -
Universal Length Generalization with Turing Programs
por: Hou, Kaiying, et al.
Publicado: (2024) -
The Role of Sparsity for Length Generalization in Transformers
por: Golowich, Noah, et al.
Publicado: (2025) -
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
por: Prabhakar, Akshara, et al.
Publicado: (2024)