Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Firdoussi, Aymane El, Seddik, Mohamed El Amine, Hayou, Soufiane, Alami, Reda, Alzubaidi, Ahmed, Hacid, Hakim |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High-dimensional Learning with Noisy Labels
by: Firdoussi, Aymane El, et al.
Published: (2024)
by: Firdoussi, Aymane El, et al.
Published: (2024)
$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling
by: Firdoussi, Aymane El, et al.
Published: (2025)
by: Firdoussi, Aymane El, et al.
Published: (2025)
Alignment with Preference Optimization Is All You Need for LLM Safety
by: Alami, Reda, et al.
Published: (2024)
by: Alami, Reda, et al.
Published: (2024)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models
by: Seddik, Mohamed El Amine
Published: (2026)
by: Seddik, Mohamed El Amine
Published: (2026)
On the Stability of the Jacobian Matrix in Deep Neural Networks
by: Dadoun, Benjamin, et al.
Published: (2025)
by: Dadoun, Benjamin, et al.
Published: (2025)
Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
by: Falcon LLM Team, et al.
Published: (2026)
by: Falcon LLM Team, et al.
Published: (2026)
A Proof of Learning Rate Transfer under $μ$P
by: Hayou, Soufiane
Published: (2025)
by: Hayou, Soufiane
Published: (2025)
Multilevel Surrogate-based Control Variates
by: Amri, Mohamed Reda El, et al.
Published: (2023)
by: Amri, Mohamed Reda El, et al.
Published: (2023)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
by: Younsi, Adam, et al.
Published: (2025)
by: Younsi, Adam, et al.
Published: (2025)
On Goodhart's law, with an application to value alignment
by: El-Mhamdi, El-Mahdi, et al.
Published: (2024)
by: El-Mhamdi, El-Mahdi, et al.
Published: (2024)
Investigating Regularization of Self-Play Language Models
by: Alami, Reda, et al.
Published: (2024)
by: Alami, Reda, et al.
Published: (2024)
A New Regression Model for Analyzing Non-Stationary Extremes in Response and Covariate Variables with an Application in Meteorology
by: Bernoussi, Amina El, et al.
Published: (2025)
by: Bernoussi, Amina El, et al.
Published: (2025)
Data Quality in Edge Machine Learning: A State-of-the-Art Survey
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2024)
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2024)
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
by: Hayou, Soufiane, et al.
Published: (2025)
by: Hayou, Soufiane, et al.
Published: (2025)
Application of Random Matrix Theory in High-Dimensional Statistics
by: Bhattacharyya, Swapnaneel, et al.
Published: (2024)
by: Bhattacharyya, Swapnaneel, et al.
Published: (2024)
A Random Matrix Theory of Pauli Tomography
by: Keenan, Nathan, et al.
Published: (2025)
by: Keenan, Nathan, et al.
Published: (2025)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026)
by: Bareilles, Gilles, et al.
Published: (2026)
Optimal sub-Gaussian variance proxy for 3-mass distributions
by: Atouani, Soufiane, et al.
Published: (2025)
by: Atouani, Soufiane, et al.
Published: (2025)
Multiscale Asymptotic Normality in Quantile Regression: Hilbert Matrices and Polynomial Designs
by: Maanan, Saïd, et al.
Published: (2025)
by: Maanan, Saïd, et al.
Published: (2025)
Maximal Inequalities for Independent Random Vectors
by: Basu, Supratik, et al.
Published: (2025)
by: Basu, Supratik, et al.
Published: (2025)
The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Generalization in Representation Models via Random Matrix Theory: Application to Recurrent Networks
by: Moakher, Yessin, et al.
Published: (2025)
by: Moakher, Yessin, et al.
Published: (2025)
A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities
by: Dandi, Yatin, et al.
Published: (2024)
by: Dandi, Yatin, et al.
Published: (2024)
Random Matrix Theory of Early-Stopped Gradient Flow: A Transient BBP Scenario
by: Coeurdoux, Florentin, et al.
Published: (2026)
by: Coeurdoux, Florentin, et al.
Published: (2026)
Machine Learning and Deep Learning in Computational Finance: A Systematic Review
by: Alami, Soufiane El Amine El, et al.
Published: (2025)
by: Alami, Soufiane El Amine El, et al.
Published: (2025)
Bootstrap inference for linear regression between variables that are never jointly observed: application in in vivo experiments
by: Arsenteva, Polina, et al.
Published: (2024)
by: Arsenteva, Polina, et al.
Published: (2024)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
by: Mansouri, Omar El, et al.
Published: (2025)
by: Mansouri, Omar El, et al.
Published: (2025)
Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning
by: Chen, Nan, et al.
Published: (2026)
by: Chen, Nan, et al.
Published: (2026)
The Impact of Initialization on LoRA Finetuning Dynamics
by: Hayou, Soufiane, et al.
Published: (2024)
by: Hayou, Soufiane, et al.
Published: (2024)
LoRA+: Efficient Low Rank Adaptation of Large Models
by: Hayou, Soufiane, et al.
Published: (2024)
by: Hayou, Soufiane, et al.
Published: (2024)
The Strong, Weak and Benign Goodhart's law. An independence-free and paradigm-agnostic formalisation
by: Majka, Adrien, et al.
Published: (2025)
by: Majka, Adrien, et al.
Published: (2025)
The Gaussian entropy map in valued fields
by: Maazouz, Yassine El
Published: (2021)
by: Maazouz, Yassine El
Published: (2021)
Goodness-of-fit testing of the distribution of posterior classification probabilities for validating model-based clustering
by: Kolei, Salima El, et al.
Published: (2025)
by: Kolei, Salima El, et al.
Published: (2025)
$σ$-Maximal Ancestral Graphs
by: Yao, Binghua, et al.
Published: (2025)
by: Yao, Binghua, et al.
Published: (2025)
Meta-Learning and representation learner: A short theoretical note
by: Bouchattaoui, Mouad El
Published: (2024)
by: Bouchattaoui, Mouad El
Published: (2024)
Regularization Using Synthetic Data in High-Dimensional Models
by: Li, Weihao, et al.
Published: (2024)
by: Li, Weihao, et al.
Published: (2024)
Random Multiplexing
by: Liu, Lei, et al.
Published: (2025)
by: Liu, Lei, et al.
Published: (2025)
Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2025)
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2025)
A Martingale Approach To Fluctuations of Rank Estimators in Sensitivity Analysis
by: Chhaibi, Reda, et al.
Published: (2026)
by: Chhaibi, Reda, et al.
Published: (2026)
Similar Items
-
High-dimensional Learning with Noisy Labels
by: Firdoussi, Aymane El, et al.
Published: (2024) -
$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling
by: Firdoussi, Aymane El, et al.
Published: (2025) -
Alignment with Preference Optimization Is All You Need for LLM Safety
by: Alami, Reda, et al.
Published: (2024) -
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024) -
How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models
by: Seddik, Mohamed El Amine
Published: (2026)