On the Generalization Ability of Unsupervised Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Yuyang, Hong, Junyuan, Zhou, Jiayu, Mahdavi, Mehrdad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Deep Gradient Leakage via Inversion Influence Functions
von: Zhang, Haobo, et al.
Veröffentlicht: (2023)
von: Zhang, Haobo, et al.
Veröffentlicht: (2023)
On the Convergence and Stability of Distributed Sub-model Training
von: Deng, Yuyang, et al.
Veröffentlicht: (2025)
von: Deng, Yuyang, et al.
Veröffentlicht: (2025)
Stochastic Compositional Minimax Optimization with Provable Convergence Guarantees
von: Deng, Yuyang, et al.
Veröffentlicht: (2024)
von: Deng, Yuyang, et al.
Veröffentlicht: (2024)
Merge before Forget: A Single LoRA Continual Learning via Continual Merging
von: Qiao, Fuli, et al.
Veröffentlicht: (2025)
von: Qiao, Fuli, et al.
Veröffentlicht: (2025)
Low-rank Momentum Factorization for Memory Efficient Training
von: Mahdavinia, Pouria, et al.
Veröffentlicht: (2025)
von: Mahdavinia, Pouria, et al.
Veröffentlicht: (2025)
Harnessing Optimization Dynamics for Curvature-Informed Model Merging
von: Mahdavinia, Pouria, et al.
Veröffentlicht: (2025)
von: Mahdavinia, Pouria, et al.
Veröffentlicht: (2025)
Model Merging via Multi-Teacher Knowledge Distillation
von: Dalili, Seyed Arshan, et al.
Veröffentlicht: (2025)
von: Dalili, Seyed Arshan, et al.
Veröffentlicht: (2025)
On the Generalization Capability of Temporal Graph Learning Algorithms: Theoretical Insights and a Simpler Method
von: Cong, Weilin, et al.
Veröffentlicht: (2024)
von: Cong, Weilin, et al.
Veröffentlicht: (2024)
Unsupervised Abnormal Stop Detection for Long Distance Coaches with Low-Frequency GPS
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
Quantum Speedups for Markov Chain Monte Carlo Methods with Application to Optimization
von: Ozgul, Guneykan, et al.
Veröffentlicht: (2025)
von: Ozgul, Guneykan, et al.
Veröffentlicht: (2025)
FedNoisy: Federated Noisy Label Learning Benchmark
von: Liang, Siqi, et al.
Veröffentlicht: (2023)
von: Liang, Siqi, et al.
Veröffentlicht: (2023)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
Discovering Bias in Latent Space: An Unsupervised Debiasing Approach
von: Adila, Dyah, et al.
Veröffentlicht: (2024)
von: Adila, Dyah, et al.
Veröffentlicht: (2024)
Shake to Leak: Fine-tuning Diffusion Models Can Amplify the Generative Privacy Risk
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
Data-Efficient Operator Learning via Unsupervised Pretraining and In-Context Learning
von: Chen, Wuyang, et al.
Veröffentlicht: (2024)
von: Chen, Wuyang, et al.
Veröffentlicht: (2024)
Unsupervised Pretraining for Fact Verification by Language Model Distillation
von: Bazaga, Adrián, et al.
Veröffentlicht: (2023)
von: Bazaga, Adrián, et al.
Veröffentlicht: (2023)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
von: Thrampoulidis, Christos, et al.
Veröffentlicht: (2025)
von: Thrampoulidis, Christos, et al.
Veröffentlicht: (2025)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
Simplicity is Key: An Unsupervised Pretraining Approach for Sparse Radio Channels
von: Ott, Jonathan, et al.
Veröffentlicht: (2025)
von: Ott, Jonathan, et al.
Veröffentlicht: (2025)
DeepOSets: Non-Autoregressive In-Context Learning with Permutation-Invariance Inductive Bias
von: Chiu, Shao-Ting, et al.
Veröffentlicht: (2024)
von: Chiu, Shao-Ting, et al.
Veröffentlicht: (2024)
Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
von: Li, Ziyue, et al.
Veröffentlicht: (2025)
von: Li, Ziyue, et al.
Veröffentlicht: (2025)
DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability
von: Cheuk, Kin Wai, et al.
Veröffentlicht: (2022)
von: Cheuk, Kin Wai, et al.
Veröffentlicht: (2022)
Mixed-Sample SGD: an End-to-end Analysis of Supervised Transfer Learning
von: Deng, Yuyang, et al.
Veröffentlicht: (2025)
von: Deng, Yuyang, et al.
Veröffentlicht: (2025)
Collaborative Learning with Different Labeling Functions
von: Deng, Yuyang, et al.
Veröffentlicht: (2024)
von: Deng, Yuyang, et al.
Veröffentlicht: (2024)
ITGPT: Generative Pretraining on Irregular Timeseries
von: Honoré, Antoine, et al.
Veröffentlicht: (2026)
von: Honoré, Antoine, et al.
Veröffentlicht: (2026)
Evaluating Loss Functions for Graph Neural Networks: Towards Pretraining and Generalization
von: Abbas, Khushnood, et al.
Veröffentlicht: (2025)
von: Abbas, Khushnood, et al.
Veröffentlicht: (2025)
Stochastic Two Points Method for Deep Model Zeroth-order Optimization
von: Pang, Yijiang, et al.
Veröffentlicht: (2024)
von: Pang, Yijiang, et al.
Veröffentlicht: (2024)
Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting
von: Moghimi, Mehrdad, et al.
Veröffentlicht: (2026)
von: Moghimi, Mehrdad, et al.
Veröffentlicht: (2026)
Exploring Graph-Transformer Out-of-Distribution Generalization Abilities
von: Niv, Itay, et al.
Veröffentlicht: (2025)
von: Niv, Itay, et al.
Veröffentlicht: (2025)
Shock-Aware Physics-Guided Fusion-DeepONet Operator for Rarefied Micro-Nozzle Flows
von: Roohi, Ehsan, et al.
Veröffentlicht: (2025)
von: Roohi, Ehsan, et al.
Veröffentlicht: (2025)
Lightweight Unsupervised Federated Learning with Pretrained Vision Language Model
von: Yan, Hao, et al.
Veröffentlicht: (2024)
von: Yan, Hao, et al.
Veröffentlicht: (2024)
Strategies for Pretraining Neural Operators
von: Zhou, Anthony, et al.
Veröffentlicht: (2024)
von: Zhou, Anthony, et al.
Veröffentlicht: (2024)
General Intelligence Requires Reward-based Pretraining
von: Han, Seungwook, et al.
Veröffentlicht: (2025)
von: Han, Seungwook, et al.
Veröffentlicht: (2025)
Why Fine-grained Labels in Pretraining Benefit Generalization?
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2024)
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2024)
Memorization Capacity of Multi-Head Attention in Transformers
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2025)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2025)
Generative Pretrained Hierarchical Transformer for Time Series Forecasting
von: Liu, Zhiding, et al.
Veröffentlicht: (2024)
von: Liu, Zhiding, et al.
Veröffentlicht: (2024)
Zebra: In-Context Generative Pretraining for Solving Parametric PDEs
von: Serrano, Louis, et al.
Veröffentlicht: (2024)
von: Serrano, Louis, et al.
Veröffentlicht: (2024)
Pretraining Scaling Laws for Generative Evaluations of Language Models
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Safety Pretraining: Toward the Next Generation of Safe AI
von: Maini, Pratyush, et al.
Veröffentlicht: (2025)
von: Maini, Pratyush, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding Deep Gradient Leakage via Inversion Influence Functions
von: Zhang, Haobo, et al.
Veröffentlicht: (2023) -
On the Convergence and Stability of Distributed Sub-model Training
von: Deng, Yuyang, et al.
Veröffentlicht: (2025) -
Stochastic Compositional Minimax Optimization with Provable Convergence Guarantees
von: Deng, Yuyang, et al.
Veröffentlicht: (2024) -
Merge before Forget: A Single LoRA Continual Learning via Continual Merging
von: Qiao, Fuli, et al.
Veröffentlicht: (2025) -
Low-rank Momentum Factorization for Memory Efficient Training
von: Mahdavinia, Pouria, et al.
Veröffentlicht: (2025)