High-dimensional Analysis of Synthetic Data Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Rezaei, Parham, Kovacevic, Filip, Locatello, Francesco, Mondelli, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery
by: Kovačević, Filip, et al.
Published: (2025)
by: Kovačević, Filip, et al.
Published: (2025)
Towards a holistic understanding of Selection Bias for Causal Effect Identification
by: Qiu, Yiwen, et al.
Published: (2026)
by: Qiu, Yiwen, et al.
Published: (2026)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
by: Kovačević, Filip, et al.
Published: (2026)
by: Kovačević, Filip, et al.
Published: (2026)
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
by: Anguita, Nicolas, et al.
Published: (2026)
by: Anguita, Nicolas, et al.
Published: (2026)
How Spurious Features Are Memorized: Precise Analysis for Random and NTK Features
by: Bombari, Simone, et al.
Published: (2023)
by: Bombari, Simone, et al.
Published: (2023)
Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization
by: Bombari, Simone, et al.
Published: (2025)
by: Bombari, Simone, et al.
Published: (2025)
High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
Causal Learning with the Invariance Principle
by: Montagna, Francesco, et al.
Published: (2026)
by: Montagna, Francesco, et al.
Published: (2026)
Improved Convergence of Score-Based Diffusion Models via Prediction-Correction
by: Pedrotti, Francesco, et al.
Published: (2023)
by: Pedrotti, Francesco, et al.
Published: (2023)
Controlling Transient Amplification Improves Long-horizon Rollouts
by: Pervez, Adeel, et al.
Published: (2026)
by: Pervez, Adeel, et al.
Published: (2026)
Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
by: Mencattini, Tommaso, et al.
Published: (2026)
by: Mencattini, Tommaso, et al.
Published: (2026)
Out-of-Distribution Detection with Relative Angles
by: Demirel, Berker, et al.
Published: (2024)
by: Demirel, Berker, et al.
Published: (2024)
Navigating the Latent Space Dynamics of Neural Models
by: Fumero, Marco, et al.
Published: (2025)
by: Fumero, Marco, et al.
Published: (2025)
Statistical and structural identifiability in representation learning
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
Privacy for Free in the Overparameterized Regime
by: Bombari, Simone, et al.
Published: (2024)
by: Bombari, Simone, et al.
Published: (2024)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
by: Bombari, Simone, et al.
Published: (2024)
by: Bombari, Simone, et al.
Published: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
High-Dimensional Private Linear Regression with Optimal Rates
by: Bombari, Simone, et al.
Published: (2025)
by: Bombari, Simone, et al.
Published: (2025)
A Law of Data Reconstruction for Random Features (and Beyond)
by: Iurada, Leonardo, et al.
Published: (2025)
by: Iurada, Leonardo, et al.
Published: (2025)
Latent Functional Maps: a spectral framework for representation alignment
by: Fumero, Marco, et al.
Published: (2024)
by: Fumero, Marco, et al.
Published: (2024)
Be More Diverse than the Most Diverse: Optimal Mixtures of Generative Models via Mixture-UCB Bandit Algorithms
by: Rezaei, Parham, et al.
Published: (2024)
by: Rezaei, Parham, et al.
Published: (2024)
Mechanistic PDE Networks for Discovery of Governing Equations
by: Pervez, Adeel, et al.
Published: (2025)
by: Pervez, Adeel, et al.
Published: (2025)
Addressing Instrument-Outcome Confounding in Mendelian Randomization through Representation Learning
by: Huang, Shimeng, et al.
Published: (2026)
by: Huang, Shimeng, et al.
Published: (2026)
Toward Identifiable Sparse Autoencoders
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
Marrying Causal Representation Learning with Dynamical Systems for Science
by: Yao, Dingling, et al.
Published: (2024)
by: Yao, Dingling, et al.
Published: (2024)
Learning Discrete Diffusion of Graphs via Free-Energy Gradient Flows
by: Rancati, Dario, et al.
Published: (2026)
by: Rancati, Dario, et al.
Published: (2026)
Optimal Regularization for Performative Learning
by: Cyffers, Edwige, et al.
Published: (2025)
by: Cyffers, Edwige, et al.
Published: (2025)
Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and Depth
by: Kögler, Kevin, et al.
Published: (2024)
by: Kögler, Kevin, et al.
Published: (2024)
Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods
by: Zhang, Yihan, et al.
Published: (2024)
by: Zhang, Yihan, et al.
Published: (2024)
Latent Space Translation via Inverse Relative Projection
by: Maiorca, Valentino, et al.
Published: (2024)
by: Maiorca, Valentino, et al.
Published: (2024)
Unifying Causal Representation Learning with the Invariance Principle
by: Yao, Dingling, et al.
Published: (2024)
by: Yao, Dingling, et al.
Published: (2024)
The Third Pillar of Causal Analysis? A Measurement Perspective on Causal Representations
by: Yao, Dingling, et al.
Published: (2025)
by: Yao, Dingling, et al.
Published: (2025)
Exploratory Causal Inference in SAEnce
by: Mencattini, Tommaso, et al.
Published: (2025)
by: Mencattini, Tommaso, et al.
Published: (2025)
Demystifying amortized causal discovery with transformers
by: Montagna, Francesco, et al.
Published: (2024)
by: Montagna, Francesco, et al.
Published: (2024)
Learning Pareto manifolds in high dimensions: How can regularization help?
by: Wegel, Tobias, et al.
Published: (2025)
by: Wegel, Tobias, et al.
Published: (2025)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
by: Súkeník, Peter, et al.
Published: (2024)
by: Súkeník, Peter, et al.
Published: (2024)
MorphGen: Controllable and Morphologically Plausible Generative Cell-Imaging
by: Demirel, Berker, et al.
Published: (2025)
by: Demirel, Berker, et al.
Published: (2025)
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
by: Wu, Diyuan, et al.
Published: (2026)
by: Wu, Diyuan, et al.
Published: (2026)
Similar Items
-
Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery
by: Kovačević, Filip, et al.
Published: (2025) -
Towards a holistic understanding of Selection Bias for Causal Effect Identification
by: Qiu, Yiwen, et al.
Published: (2026) -
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
by: Kovačević, Filip, et al.
Published: (2026) -
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
by: Anguita, Nicolas, et al.
Published: (2026) -
How Spurious Features Are Memorized: Precise Analysis for Random and NTK Features
by: Bombari, Simone, et al.
Published: (2023)