Towards Understanding Generalization in DP-GD: A Case Study in Training Two-Layer CNNs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Zhongjie, Wang, Puyu, Zhang, Chenyang, Cao, Yuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding the Benefits of SimCLR Pre-Training in Two-Layer Convolutional Neural Networks
von: Zhang, Han, et al.
Veröffentlicht: (2024)
von: Zhang, Han, et al.
Veröffentlicht: (2024)
Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise
von: Wang, Puyu, et al.
Veröffentlicht: (2026)
von: Wang, Puyu, et al.
Veröffentlicht: (2026)
Looped Transformers with Layer Normalization Provably Learn the Power Method
von: Wu, Lyumin, et al.
Veröffentlicht: (2026)
von: Wu, Lyumin, et al.
Veröffentlicht: (2026)
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
von: Wang, Puyu, et al.
Veröffentlicht: (2023)
von: Wang, Puyu, et al.
Veröffentlicht: (2023)
Towards Understanding Transformers in Learning Random Walks
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Planning with Language and Generative Models: Toward General Reward-Guided Wireless Network Design
von: Yuan, Chenyang, et al.
Veröffentlicht: (2026)
von: Yuan, Chenyang, et al.
Veröffentlicht: (2026)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
Generalization analysis with deep ReLU networks for metric and similarity learning
von: Zhou, Junyu, et al.
Veröffentlicht: (2024)
von: Zhou, Junyu, et al.
Veröffentlicht: (2024)
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
Transformers Trained via Gradient Descent Can Provably Learn a Class of Teacher Models
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
FlashDP: Private Training Large Language Models with Efficient DP-SGD
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
Understanding CNNs from excitations
von: Ying, Zijian, et al.
Veröffentlicht: (2022)
von: Ying, Zijian, et al.
Veröffentlicht: (2022)
Can overfitted deep neural networks in adversarial training generalize? -- An approximation viewpoint
von: Shi, Zhongjie, et al.
Veröffentlicht: (2024)
von: Shi, Zhongjie, et al.
Veröffentlicht: (2024)
Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
von: Pavlitska, Svetlana, et al.
Veröffentlicht: (2025)
von: Pavlitska, Svetlana, et al.
Veröffentlicht: (2025)
Differential Privacy in Two-Layer Networks: How DP-SGD Harms Fairness and Robustness
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
von: Shi, Zhongjie, et al.
Veröffentlicht: (2026)
von: Shi, Zhongjie, et al.
Veröffentlicht: (2026)
Initialization Matters: On the Benign Overfitting of Two-Layer ReLU CNN with Fully Trainable Layers
von: Shang, Shuning, et al.
Veröffentlicht: (2024)
von: Shang, Shuning, et al.
Veröffentlicht: (2024)
Transformer Learns Optimal Variable Selection in Group-Sparse Classification
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
The Implicit Bias of Adam on Separable Data
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov-Arnold Networks
von: Wang, Puyu, et al.
Veröffentlicht: (2026)
von: Wang, Puyu, et al.
Veröffentlicht: (2026)
Fourier Circuits in Neural Networks and Transformers: A Case Study of Modular Arithmetic with Multiple Inputs
von: Li, Chenyang, et al.
Veröffentlicht: (2024)
von: Li, Chenyang, et al.
Veröffentlicht: (2024)
Towards Theoretical Understandings of Self-Consuming Generative Models
von: Fu, Shi, et al.
Veröffentlicht: (2024)
von: Fu, Shi, et al.
Veröffentlicht: (2024)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
von: Chen, Lu, et al.
Veröffentlicht: (2025)
von: Chen, Lu, et al.
Veröffentlicht: (2025)
R+R:Understanding Hyperparameter Effects in DP-SGD
von: Morsbach, Felix, et al.
Veröffentlicht: (2024)
von: Morsbach, Felix, et al.
Veröffentlicht: (2024)
OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
DP-MicroAdam: Private and Frugal Algorithm for Training and Fine-tuning
von: Hudişteanu, Mihaela, et al.
Veröffentlicht: (2025)
von: Hudişteanu, Mihaela, et al.
Veröffentlicht: (2025)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
von: Bombari, Simone, et al.
Veröffentlicht: (2024)
von: Bombari, Simone, et al.
Veröffentlicht: (2024)
DP-DGAD: A Generalist Dynamic Graph Anomaly Detector with Dynamic Prototypes
von: Zheng, Jialun, et al.
Veröffentlicht: (2025)
von: Zheng, Jialun, et al.
Veröffentlicht: (2025)
Training a Two Layer ReLU Network Analytically
von: Barbu, Adrian
Veröffentlicht: (2023)
von: Barbu, Adrian
Veröffentlicht: (2023)
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2025)
DP-λCGD: Efficient Noise Correlation for Differentially Private Model Training
von: Kalinin, Nikita P., et al.
Veröffentlicht: (2026)
von: Kalinin, Nikita P., et al.
Veröffentlicht: (2026)
Theory of Decentralized Robust Kernel-Based Learning
von: Yu, Zhan, et al.
Veröffentlicht: (2025)
von: Yu, Zhan, et al.
Veröffentlicht: (2025)
Constant Stepsize Local GD for Logistic Regression: Acceleration by Instability
von: Crawshaw, Michael, et al.
Veröffentlicht: (2025)
von: Crawshaw, Michael, et al.
Veröffentlicht: (2025)
On the Stability of Nonlinear Dynamics in GD and SGD: Beyond Quadratic Potentials
von: Mulayoff, Rotem, et al.
Veröffentlicht: (2026)
von: Mulayoff, Rotem, et al.
Veröffentlicht: (2026)
ManifoldGD: Training-Free Hierarchical Manifold Guidance for Diffusion-Based Dataset Distillation
von: Roy, Ayush, et al.
Veröffentlicht: (2026)
von: Roy, Ayush, et al.
Veröffentlicht: (2026)
Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning
von: Alam, Md Tanvirul, et al.
Veröffentlicht: (2025)
von: Alam, Md Tanvirul, et al.
Veröffentlicht: (2025)
Integrating Feature Correlation in Differential Privacy with Applications in DP-ERM
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
Scalable Lipschitz Estimation for CNNs
von: Sulehman, Yusuf, et al.
Veröffentlicht: (2024)
von: Sulehman, Yusuf, et al.
Veröffentlicht: (2024)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
von: Kwon, Sangwoo, et al.
Veröffentlicht: (2025)
von: Kwon, Sangwoo, et al.
Veröffentlicht: (2025)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding the Benefits of SimCLR Pre-Training in Two-Layer Convolutional Neural Networks
von: Zhang, Han, et al.
Veröffentlicht: (2024) -
Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise
von: Wang, Puyu, et al.
Veröffentlicht: (2026) -
Looped Transformers with Layer Normalization Provably Learn the Power Method
von: Wu, Lyumin, et al.
Veröffentlicht: (2026) -
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
von: Wang, Puyu, et al.
Veröffentlicht: (2023) -
Towards Understanding Transformers in Learning Random Walks
von: Shi, Wei, et al.
Veröffentlicht: (2025)