Scale Efficient Training for Large Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Qing, Gao, Junyu, Wang, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Large-Scale Dataset Distillation via Exploration-Exploitation Optimization
by: Alahmadi, Muhammad J., et al.
Published: (2026)
by: Alahmadi, Muhammad J., et al.
Published: (2026)
Aneumo: A Large-Scale Comprehensive Synthetic Dataset of Aneurysm Hemodynamics
by: Li, Xigui, et al.
Published: (2025)
by: Li, Xigui, et al.
Published: (2025)
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
by: Kišš, Martin, et al.
Published: (2025)
by: Kišš, Martin, et al.
Published: (2025)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
Kaputt: A Large-Scale Dataset for Visual Defect Detection
by: Höfer, Sebastian, et al.
Published: (2025)
by: Höfer, Sebastian, et al.
Published: (2025)
Soft Label Pruning and Quantization for Large-Scale Dataset Distillation
by: Lingao, Xiao, et al.
Published: (2026)
by: Lingao, Xiao, et al.
Published: (2026)
E$^{2}$GAN: Efficient Training of Efficient GANs for Image-to-Image Translation
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving
by: Embacher, Felix, et al.
Published: (2026)
by: Embacher, Felix, et al.
Published: (2026)
Summer-22B: A Systematic Approach to Dataset Engineering and Training at Scale for Video Foundation Model
by: Ryu, Simo, et al.
Published: (2026)
by: Ryu, Simo, et al.
Published: (2026)
Efficient Long-Horizon GUI Agents via Training-Free KV Cache Compression
by: Zhou, Bowen, et al.
Published: (2026)
by: Zhou, Bowen, et al.
Published: (2026)
DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi Text
by: Singh, Kunwar Arpit, et al.
Published: (2026)
by: Singh, Kunwar Arpit, et al.
Published: (2026)
MaskOpt: A Large-Scale Mask Optimization Dataset to Advance AI in Integrated Circuit Manufacturing
by: Hu, Yuting, et al.
Published: (2025)
by: Hu, Yuting, et al.
Published: (2025)
Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection
by: Zhang, Yaoteng, et al.
Published: (2026)
by: Zhang, Yaoteng, et al.
Published: (2026)
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
by: Li, Yitong, et al.
Published: (2026)
by: Li, Yitong, et al.
Published: (2026)
On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm
by: Sun, Peng, et al.
Published: (2023)
by: Sun, Peng, et al.
Published: (2023)
Efficient Training with Denoised Neural Weights
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
Architecture, Dataset and Model-Scale Agnostic Data-free Meta-Learning
by: Hu, Zixuan, et al.
Published: (2023)
by: Hu, Zixuan, et al.
Published: (2023)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
COMO: Closed-Loop Optical Molecule Recognition with Minimum Risk Training
by: Lyu, Zhuoqi, et al.
Published: (2026)
by: Lyu, Zhuoqi, et al.
Published: (2026)
Video Understanding by Design: How Datasets Shape Architectures and Insights
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
Quantum-inspired Interpretable Deep Learning Architecture for Text Sentiment Analysis
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
Scaling Diffusion Transformers Efficiently via $μ$P
by: Zheng, Chenyu, et al.
Published: (2025)
by: Zheng, Chenyu, et al.
Published: (2025)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
by: Zhou, Gengze, et al.
Published: (2025)
by: Zhou, Gengze, et al.
Published: (2025)
Conditioning GAN Without Training Dataset
by: Mekonnen, Kidist Amde
Published: (2024)
by: Mekonnen, Kidist Amde
Published: (2024)
COAP: Memory-Efficient Training with Correlation-Aware Gradient Projection
by: Xiao, Jinqi, et al.
Published: (2024)
by: Xiao, Jinqi, et al.
Published: (2024)
DELT: A Simple Diversity-driven EarlyLate Training for Dataset Distillation
by: Shen, Zhiqiang, et al.
Published: (2024)
by: Shen, Zhiqiang, et al.
Published: (2024)
EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
by: VanVoorst, Brian, et al.
Published: (2026)
by: VanVoorst, Brian, et al.
Published: (2026)
Multi-Level Feature Distillation of Joint Teachers Trained on Distinct Image Datasets
by: Iordache, Adrian, et al.
Published: (2024)
by: Iordache, Adrian, et al.
Published: (2024)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
by: Nezhurina, Marianna, et al.
Published: (2025)
by: Nezhurina, Marianna, et al.
Published: (2025)
DeiT-LT Distillation Strikes Back for Vision Transformer Training on Long-Tailed Datasets
by: Rangwani, Harsh, et al.
Published: (2024)
by: Rangwani, Harsh, et al.
Published: (2024)
Adaptive Training Meets Progressive Scaling: Elevating Efficiency in Diffusion Models
by: Li, Wenhao, et al.
Published: (2023)
by: Li, Wenhao, et al.
Published: (2023)
New York Smells: A Large Multimodal Dataset for Olfaction
by: Ozguroglu, Ege, et al.
Published: (2025)
by: Ozguroglu, Ege, et al.
Published: (2025)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
by: Tang, Haotian, et al.
Published: (2024)
by: Tang, Haotian, et al.
Published: (2024)
Like Humans to Few-Shot Learning through Knowledge Permeation of Vision and Text
by: Jia, Yuyu, et al.
Published: (2024)
by: Jia, Yuyu, et al.
Published: (2024)
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
by: Jimenez-Perez, Guillermo, et al.
Published: (2024)
by: Jimenez-Perez, Guillermo, et al.
Published: (2024)
Elucidating the Design Space of Dataset Condensation
by: Shao, Shitong, et al.
Published: (2024)
by: Shao, Shitong, et al.
Published: (2024)
Similar Items
-
Accelerating Large-Scale Dataset Distillation via Exploration-Exploitation Optimization
by: Alahmadi, Muhammad J., et al.
Published: (2026) -
Aneumo: A Large-Scale Comprehensive Synthetic Dataset of Aneurysm Hemodynamics
by: Li, Xigui, et al.
Published: (2025) -
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
by: Kišš, Martin, et al.
Published: (2025) -
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024) -
Kaputt: A Large-Scale Dataset for Visual Defect Detection
by: Höfer, Sebastian, et al.
Published: (2025)