On Pretraining Data Diversity for Self-Supervised Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hammoud, Hasan Abed Al Kader, Das, Tuhin, Pizzati, Fabio, Torr, Philip, Bibi, Adel, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
On the Importance of Pretraining Data Alignment for Atomic Property Prediction
by: Ghunaim, Yasir, et al.
Published: (2025)
by: Ghunaim, Yasir, et al.
Published: (2025)
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
MatchDiffusion: Training-free Generation of Match-cuts
by: Pardo, Alejandro, et al.
Published: (2024)
by: Pardo, Alejandro, et al.
Published: (2024)
Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
by: Zhang, Wenxuan, et al.
Published: (2024)
by: Zhang, Wenxuan, et al.
Published: (2024)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
by: Iakovleva, Ekaterina, et al.
Published: (2024)
by: Iakovleva, Ekaterina, et al.
Published: (2024)
FedMedICL: Towards Holistic Evaluation of Distribution Shifts in Federated Medical Imaging
by: Alhamoud, Kumail, et al.
Published: (2024)
by: Alhamoud, Kumail, et al.
Published: (2024)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024)
by: Pondaven, Alexander, et al.
Published: (2024)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
by: Prabhu, Ameya, et al.
Published: (2023)
by: Prabhu, Ameya, et al.
Published: (2023)
On the Coexistence and Ensembling of Watermarks
by: Petrov, Aleksandar, et al.
Published: (2025)
by: Petrov, Aleksandar, et al.
Published: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
GenView: Enhancing View Quality with Pretrained Generative Model for Self-Supervised Learning
by: Li, Xiaojie, et al.
Published: (2024)
by: Li, Xiaojie, et al.
Published: (2024)
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
by: Eiras, Francisco, et al.
Published: (2023)
by: Eiras, Francisco, et al.
Published: (2023)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
ActionParty: Multi-Subject Action Binding in Generative Video Games
by: Pondaven, Alexander, et al.
Published: (2026)
by: Pondaven, Alexander, et al.
Published: (2026)
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
by: Yuan, Jianhao, et al.
Published: (2025)
by: Yuan, Jianhao, et al.
Published: (2025)
SimCS: Simulation for Domain Incremental Online Continual Segmentation
by: Alfarra, Motasem, et al.
Published: (2022)
by: Alfarra, Motasem, et al.
Published: (2022)
Learning Semantic Segmentation with Query Points Supervision on Aerial Images
by: Rivier, Santiago, et al.
Published: (2023)
by: Rivier, Santiago, et al.
Published: (2023)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
MVTN: Learning Multi-View Transformations for 3D Understanding
by: Hamdi, Abdullah, et al.
Published: (2022)
by: Hamdi, Abdullah, et al.
Published: (2022)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
TAPS: Task Aware Proposal Distributions for Speculative Sampling
by: Zbib, Mohamad, et al.
Published: (2026)
by: Zbib, Mohamad, et al.
Published: (2026)
FactoFormer: Factorized Hyperspectral Transformers with Self-Supervised Pretraining
by: Mohamed, Shaheer, et al.
Published: (2023)
by: Mohamed, Shaheer, et al.
Published: (2023)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
MAE-Based Self-Supervised Pretraining for Data-Efficient Medical Image Segmentation Using nnFormer
by: Sureddi, R. M. Krishna, et al.
Published: (2026)
by: Sureddi, R. M. Krishna, et al.
Published: (2026)
SAR Object Detection with Self-Supervised Pretraining and Curriculum-Aware Sampling
by: Almalioglu, Yasin, et al.
Published: (2025)
by: Almalioglu, Yasin, et al.
Published: (2025)
SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval
by: Jawade, Bhavin, et al.
Published: (2025)
by: Jawade, Bhavin, et al.
Published: (2025)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
by: Khalifi, Omar El, et al.
Published: (2026)
by: Khalifi, Omar El, et al.
Published: (2026)
Benchmarking Robust Self-Supervised Learning Across Diverse Downstream Tasks
by: Kowalczuk, Antoni, et al.
Published: (2024)
by: Kowalczuk, Antoni, et al.
Published: (2024)
TinySSL: Distilled Self-Supervised Pretraining for Sub-Megabyte MCU Models
by: Wilson, Bibin
Published: (2026)
by: Wilson, Bibin
Published: (2026)
MessyKitchens: Contact-rich object-level 3D scene reconstruction
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
Mitigating Overfitting in Medical Imaging: Self-Supervised Pretraining vs. ImageNet Transfer Learning for Dermatological Diagnosis
by: Matas, Iván, et al.
Published: (2025)
by: Matas, Iván, et al.
Published: (2025)
Similar Items
-
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024) -
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025) -
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024) -
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
by: Alssum, Lama, et al.
Published: (2025) -
On the Importance of Pretraining Data Alignment for Atomic Property Prediction
by: Ghunaim, Yasir, et al.
Published: (2025)