Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Anagnostidis, Sotiris, Bachmann, Gregor, Schlag, Imanol, Hofmann, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Meta-Pruning via Optimal Transport
di: Theus, Alexander, et al.
Pubblicazione: (2024)
di: Theus, Alexander, et al.
Pubblicazione: (2024)
FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2025)
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2025)
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
di: Bill, Eric Tillman, et al.
Pubblicazione: (2025)
di: Bill, Eric Tillman, et al.
Pubblicazione: (2025)
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
di: Yang, Han, et al.
Pubblicazione: (2024)
di: Yang, Han, et al.
Pubblicazione: (2024)
Energy Scaling Laws for Diffusion Models: Quantifying Compute in Image Generation
di: Iyengar, Aniketh, et al.
Pubblicazione: (2025)
di: Iyengar, Aniketh, et al.
Pubblicazione: (2025)
IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait
di: Yang, Han, et al.
Pubblicazione: (2025)
di: Yang, Han, et al.
Pubblicazione: (2025)
Training a Computer Vision Model for Commercial Bakeries with Primarily Synthetic Images
di: Schmitt, Thomas H., et al.
Pubblicazione: (2024)
di: Schmitt, Thomas H., et al.
Pubblicazione: (2024)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2023)
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2023)
ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
di: Lu, Shunlin, et al.
Pubblicazione: (2024)
di: Lu, Shunlin, et al.
Pubblicazione: (2024)
Adaptive Training Meets Progressive Scaling: Elevating Efficiency in Diffusion Models
di: Li, Wenhao, et al.
Pubblicazione: (2023)
di: Li, Wenhao, et al.
Pubblicazione: (2023)
Variance-Aware Adaptive Weighting for Diffusion Model Training
di: Sun, Nanlong, et al.
Pubblicazione: (2026)
di: Sun, Nanlong, et al.
Pubblicazione: (2026)
Scaling Laws for Black box Adversarial Attacks
di: Liu, Chuan, et al.
Pubblicazione: (2024)
di: Liu, Chuan, et al.
Pubblicazione: (2024)
Towards Large-Scale Training of Pathology Foundation Models
di: ai, kaiko., et al.
Pubblicazione: (2024)
di: ai, kaiko., et al.
Pubblicazione: (2024)
AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
di: Miranda, Imanol, et al.
Pubblicazione: (2026)
di: Miranda, Imanol, et al.
Pubblicazione: (2026)
Class Adaptive Conformal Training
di: Marani, Badr-Eddine, et al.
Pubblicazione: (2026)
di: Marani, Badr-Eddine, et al.
Pubblicazione: (2026)
Adaptive Non-uniform Timestep Sampling for Accelerating Diffusion Model Training
di: Kim, Myunsoo, et al.
Pubblicazione: (2024)
di: Kim, Myunsoo, et al.
Pubblicazione: (2024)
Training Feature Attribution for Vision Models
di: Bacha, Aziz, et al.
Pubblicazione: (2025)
di: Bacha, Aziz, et al.
Pubblicazione: (2025)
RB-Modulation: Training-Free Personalization of Diffusion Models using Stochastic Optimal Control
di: Rout, Litu, et al.
Pubblicazione: (2024)
di: Rout, Litu, et al.
Pubblicazione: (2024)
Graph Neural Networks for Surgical Scene Segmentation
di: Li, Yihan, et al.
Pubblicazione: (2025)
di: Li, Yihan, et al.
Pubblicazione: (2025)
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers
di: Liu, Hongbo
Pubblicazione: (2024)
di: Liu, Hongbo
Pubblicazione: (2024)
Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
di: Bergmeister, Andreas, et al.
Pubblicazione: (2026)
di: Bergmeister, Andreas, et al.
Pubblicazione: (2026)
Revisiting Class-Incremental Learning with Pre-Trained Models: Generalizability and Adaptivity are All You Need
di: Zhou, Da-Wei, et al.
Pubblicazione: (2023)
di: Zhou, Da-Wei, et al.
Pubblicazione: (2023)
On Domain-Adaptive Post-Training for Multimodal Large Language Models
di: Cheng, Daixuan, et al.
Pubblicazione: (2024)
di: Cheng, Daixuan, et al.
Pubblicazione: (2024)
How much is a noisy image worth? Data Scaling Laws for Ambient Diffusion
di: Daras, Giannis, et al.
Pubblicazione: (2024)
di: Daras, Giannis, et al.
Pubblicazione: (2024)
TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce Edge
di: Kwon, Young D., et al.
Pubblicazione: (2023)
di: Kwon, Young D., et al.
Pubblicazione: (2023)
Data Scaling Laws for End-to-End Autonomous Driving
di: Naumann, Alexander, et al.
Pubblicazione: (2025)
di: Naumann, Alexander, et al.
Pubblicazione: (2025)
Adding simple structure at inference improves Vision-Language Compositionality
di: Miranda, Imanol, et al.
Pubblicazione: (2025)
di: Miranda, Imanol, et al.
Pubblicazione: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
di: Miranda, Imanol, et al.
Pubblicazione: (2024)
di: Miranda, Imanol, et al.
Pubblicazione: (2024)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
di: Chen, Haoran, et al.
Pubblicazione: (2024)
di: Chen, Haoran, et al.
Pubblicazione: (2024)
dynActivation: A Trainable Activation Family for Adaptive Nonlinearity
di: Bachmann, Alois
Pubblicazione: (2026)
di: Bachmann, Alois
Pubblicazione: (2026)
Learning Images Across Scales Using Adversarial Training
di: Wolski, Krzysztof, et al.
Pubblicazione: (2024)
di: Wolski, Krzysztof, et al.
Pubblicazione: (2024)
Wild Visual Navigation: Fast Traversability Learning via Pre-Trained Models and Online Self-Supervision
di: Mattamala, Matías, et al.
Pubblicazione: (2024)
di: Mattamala, Matías, et al.
Pubblicazione: (2024)
From Isolation to Integration: Building an Adaptive Expert Forest for Pre-Trained Model-based Class-Incremental Learning
di: Liu, Ruiqi, et al.
Pubblicazione: (2026)
di: Liu, Ruiqi, et al.
Pubblicazione: (2026)
Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream
di: Gokce, Abdulkadir, et al.
Pubblicazione: (2024)
di: Gokce, Abdulkadir, et al.
Pubblicazione: (2024)
Towards Fair Class-wise Robustness: Class Optimal Distribution Adversarial Training
di: Zhi, Hongxin, et al.
Pubblicazione: (2025)
di: Zhi, Hongxin, et al.
Pubblicazione: (2025)
Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural Networks
di: Zou, Jinping, et al.
Pubblicazione: (2024)
di: Zou, Jinping, et al.
Pubblicazione: (2024)
Taming the Long Tail: Rebalancing Adversarial Training via Adaptive Perturbation
di: Zhang, Lilin, et al.
Pubblicazione: (2026)
di: Zhang, Lilin, et al.
Pubblicazione: (2026)
Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws
di: Wei, Xiyuan, et al.
Pubblicazione: (2025)
di: Wei, Xiyuan, et al.
Pubblicazione: (2025)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
di: Ramachandran, Rahul, et al.
Pubblicazione: (2025)
di: Ramachandran, Rahul, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards Meta-Pruning via Optimal Transport
di: Theus, Alexander, et al.
Pubblicazione: (2024) -
FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2025) -
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
di: Bill, Eric Tillman, et al.
Pubblicazione: (2025) -
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
di: Yang, Han, et al.
Pubblicazione: (2024) -
Energy Scaling Laws for Diffusion Models: Quantifying Compute in Image Generation
di: Iyengar, Aniketh, et al.
Pubblicazione: (2025)