When Data Falls Short: Grokking Below the Critical Threshold
Fuente:
arXiv
Guardado en:
| Autores principales: | Singh, Vaibhav, Belilovsky, Eugene, Aljundi, Rahaf |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
por: Singh, Vaibhav, et al.
Publicado: (2024)
por: Singh, Vaibhav, et al.
Publicado: (2024)
Model Parallelism With Subnetwork Data Parallelism
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
por: Nanfack, Geraldin, et al.
Publicado: (2025)
por: Nanfack, Geraldin, et al.
Publicado: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
por: Miahi, Erfan, et al.
Publicado: (2026)
por: Miahi, Erfan, et al.
Publicado: (2026)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
por: Singh, Vaibhav, et al.
Publicado: (2025)
por: Singh, Vaibhav, et al.
Publicado: (2025)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
por: Dorovatas, Vaggelis, et al.
Publicado: (2025)
por: Dorovatas, Vaggelis, et al.
Publicado: (2025)
Model Breadcrumbs: Scaling Multi-Task Model Merging with Sparse Masks
por: Davari, MohammadReza, et al.
Publicado: (2023)
por: Davari, MohammadReza, et al.
Publicado: (2023)
Stabilizing Native Low-Rank LLM Pretraining
por: Janson, Paul, et al.
Publicado: (2026)
por: Janson, Paul, et al.
Publicado: (2026)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
por: Legate, Gwen, et al.
Publicado: (2025)
por: Legate, Gwen, et al.
Publicado: (2025)
Celo2: Towards Learned Optimization Free Lunch
por: Moudgil, Abhinav, et al.
Publicado: (2026)
por: Moudgil, Abhinav, et al.
Publicado: (2026)
Efficient Refusal Ablation in LLM through Optimal Transport
por: Nanfack, Geraldin, et al.
Publicado: (2026)
por: Nanfack, Geraldin, et al.
Publicado: (2026)
Celo: Training Versatile Learned Optimizers on a Compute Diet
por: Moudgil, Abhinav, et al.
Publicado: (2025)
por: Moudgil, Abhinav, et al.
Publicado: (2025)
Communication Efficient LLM Pre-training with SparseLoCo
por: Sarfi, Amir, et al.
Publicado: (2025)
por: Sarfi, Amir, et al.
Publicado: (2025)
To Grok Grokking: Provable Grokking in Ridge Regression
por: Xu, Mingyue, et al.
Publicado: (2026)
por: Xu, Mingyue, et al.
Publicado: (2026)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
por: Hameed, Humza Wajid, et al.
Publicado: (2024)
por: Hameed, Humza Wajid, et al.
Publicado: (2024)
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models
por: Panos, Aristeidis, et al.
Publicado: (2024)
por: Panos, Aristeidis, et al.
Publicado: (2024)
Critical Data Size of Language Models from a Grokking Perspective
por: Zhu, Xuekai, et al.
Publicado: (2024)
por: Zhu, Xuekai, et al.
Publicado: (2024)
DragD3D: Realistic Mesh Editing with Rigidity Control Driven by 2D Diffusion Priors
por: Xie, Tianhao, et al.
Publicado: (2023)
por: Xie, Tianhao, et al.
Publicado: (2023)
Rethinking Prompt Optimization: Reinforcement, Diversification, and Migration in Blackbox LLMs
por: Davari, MohammadReza, et al.
Publicado: (2025)
por: Davari, MohammadReza, et al.
Publicado: (2025)
Explaining Grokking in Transformers through the Lens of Inductive Bias
por: Singh, Jaisidh, et al.
Publicado: (2026)
por: Singh, Jaisidh, et al.
Publicado: (2026)
From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation
por: Nanfack, Geraldin, et al.
Publicado: (2024)
por: Nanfack, Geraldin, et al.
Publicado: (2024)
When Active Learning Falls Short: An Empirical Study on Chemical Reaction Extraction
por: Yu, Simin, et al.
Publicado: (2026)
por: Yu, Simin, et al.
Publicado: (2026)
Grokking as a Variance-Limited Phase Transition: Spectral Gating and the Epsilon-Stability Threshold
por: Acharya, Pratyush, et al.
Publicado: (2026)
por: Acharya, Pratyush, et al.
Publicado: (2026)
Controlling Grokking with Nonlinearity and Data Symmetry
por: Salah, Ahmed, et al.
Publicado: (2024)
por: Salah, Ahmed, et al.
Publicado: (2024)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
por: Guimard, Quentin, et al.
Publicado: (2026)
por: Guimard, Quentin, et al.
Publicado: (2026)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
por: Thérien, Benjamin, et al.
Publicado: (2025)
por: Thérien, Benjamin, et al.
Publicado: (2025)
Heterogeneous Low-Bandwidth Pre-Training of LLMs
por: Obeidi, Yazan, et al.
Publicado: (2026)
por: Obeidi, Yazan, et al.
Publicado: (2026)
The Complexity Dynamics of Grokking
por: DeMoss, Branton, et al.
Publicado: (2024)
por: DeMoss, Branton, et al.
Publicado: (2024)
Measuring Sharpness in Grokking
por: Miller, Jack, et al.
Publicado: (2024)
por: Miller, Jack, et al.
Publicado: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
por: Minegishi, Gouki, et al.
Publicado: (2023)
por: Minegishi, Gouki, et al.
Publicado: (2023)
Less is More: Undertraining Experts Improves Model Upcycling
por: Horoi, Stefan, et al.
Publicado: (2025)
por: Horoi, Stefan, et al.
Publicado: (2025)
Evaluating Imputation Techniques for Short-Term Gaps in Heart Rate Data
por: Gupta, Vaibhav, et al.
Publicado: (2025)
por: Gupta, Vaibhav, et al.
Publicado: (2025)
Meta-learning Optimizers for Communication-Efficient Learning
por: Joseph, Charles-Étienne, et al.
Publicado: (2023)
por: Joseph, Charles-Étienne, et al.
Publicado: (2023)
Grokked Models are Better Unlearners
por: Liang, Yuanbang, et al.
Publicado: (2025)
por: Liang, Yuanbang, et al.
Publicado: (2025)
Non-Uniform Parameter-Wise Model Merging
por: Camacho, Albert Manuel Orozco, et al.
Publicado: (2024)
por: Camacho, Albert Manuel Orozco, et al.
Publicado: (2024)
Topological Signatures of Grokking
por: Tang, Yifan, et al.
Publicado: (2026)
por: Tang, Yifan, et al.
Publicado: (2026)
Generalization Below the Edge of Stability: The Role of Data Geometry
por: Liang, Tongtong, et al.
Publicado: (2025)
por: Liang, Tongtong, et al.
Publicado: (2025)
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
PyLO: Towards Accessible Learned Optimizers in PyTorch
por: Janson, Paul, et al.
Publicado: (2025)
por: Janson, Paul, et al.
Publicado: (2025)
Ejemplares similares
-
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
por: Singh, Vaibhav, et al.
Publicado: (2024) -
Model Parallelism With Subnetwork Data Parallelism
por: Singh, Vaibhav, et al.
Publicado: (2025) -
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
por: Singh, Vaibhav, et al.
Publicado: (2025) -
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
por: Nanfack, Geraldin, et al.
Publicado: (2025) -
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
por: Miahi, Erfan, et al.
Publicado: (2026)