Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
Fuente:
arXiv
Saved in:
| Main Authors: | Alabdulmohsin, Ibrahim, Zhai, Xiaohua, Kolesnikov, Alexander, Beyer, Lucas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
by: Yavari, Sara, et al.
Published: (2025)
by: Yavari, Sara, et al.
Published: (2025)
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
by: Zhou, Fang, et al.
Published: (2025)
by: Zhou, Fang, et al.
Published: (2025)
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
by: Eymaël, Alexandre, et al.
Published: (2024)
by: Eymaël, Alexandre, et al.
Published: (2024)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
by: Tang, Guowei
Published: (2026)
by: Tang, Guowei
Published: (2026)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
by: Qian, Wenxu, et al.
Published: (2025)
by: Qian, Wenxu, et al.
Published: (2025)
Adaptive Self-Training for Object Detection
by: Vandeghen, Renaud, et al.
Published: (2022)
by: Vandeghen, Renaud, et al.
Published: (2022)
Pointing-Based Object Recognition
by: Hajdúch, Lukáš, et al.
Published: (2026)
by: Hajdúch, Lukáš, et al.
Published: (2026)
Contrastive pretraining for semantic segmentation is robust to noisy positive pairs
by: Gerard, Sebastian, et al.
Published: (2022)
by: Gerard, Sebastian, et al.
Published: (2022)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
by: Hou, Zhangcheng, et al.
Published: (2026)
by: Hou, Zhangcheng, et al.
Published: (2026)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026)
by: Dong, Mingkang, et al.
Published: (2026)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
by: Seo, Huichan, et al.
Published: (2025)
by: Seo, Huichan, et al.
Published: (2025)
Smooth regularization for efficient video recognition
by: Goldman, Gil, et al.
Published: (2025)
by: Goldman, Gil, et al.
Published: (2025)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
by: Bergkvist, Viktor, et al.
Published: (2026)
by: Bergkvist, Viktor, et al.
Published: (2026)
The Power of Next-Frame Prediction for Learning Physical Laws
by: Winterbottom, Thomas, et al.
Published: (2024)
by: Winterbottom, Thomas, et al.
Published: (2024)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
by: Wang, Zhaohui, et al.
Published: (2025)
by: Wang, Zhaohui, et al.
Published: (2025)
Conditional Compatibility Learning for Context-Dependent Anomaly Detection
by: Mishra, Shashank, et al.
Published: (2026)
by: Mishra, Shashank, et al.
Published: (2026)
Invariant Representation via Decoupling Style and Spurious Features from Images
by: Li, Ruimeng, et al.
Published: (2023)
by: Li, Ruimeng, et al.
Published: (2023)
Real-Time Flying Object Detection with YOLOv8
by: Reis, Dillon, et al.
Published: (2023)
by: Reis, Dillon, et al.
Published: (2023)
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
ESCAPE: Energy-based Selective Adaptive Correction for Out-of-distribution 3D Human Pose Estimation
by: Bidulka, Luke, et al.
Published: (2024)
by: Bidulka, Luke, et al.
Published: (2024)
Engineering an Efficient Object Tracker for Non-Linear Motion
by: Adžemović, Momir, et al.
Published: (2024)
by: Adžemović, Momir, et al.
Published: (2024)
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
by: Huang, Yuhang, et al.
Published: (2025)
by: Huang, Yuhang, et al.
Published: (2025)
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
by: Bu, Weijue, et al.
Published: (2025)
by: Bu, Weijue, et al.
Published: (2025)
Auto-Annotation Quality Prediction for Semi-Supervised Learning with Ensembles
by: Simon, Dror, et al.
Published: (2019)
by: Simon, Dror, et al.
Published: (2019)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
by: Mansour, Jad, et al.
Published: (2025)
by: Mansour, Jad, et al.
Published: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
by: Mansour, Jad, et al.
Published: (2024)
by: Mansour, Jad, et al.
Published: (2024)
Training a Student Expert via Semi-Supervised Foundation Model Distillation
by: Taghavi, Pardis, et al.
Published: (2026)
by: Taghavi, Pardis, et al.
Published: (2026)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
by: Boudras, Thomas, et al.
Published: (2025)
by: Boudras, Thomas, et al.
Published: (2025)
FlightScope: An Experimental Comparative Review of Aircraft Detection Algorithms in Satellite Imagery
by: Ghazouali, Safouane El, et al.
Published: (2024)
by: Ghazouali, Safouane El, et al.
Published: (2024)
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
by: Ziakas, Christos, et al.
Published: (2025)
by: Ziakas, Christos, et al.
Published: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026)
by: Liu, Zhi
Published: (2026)
NFIG: Multi-Scale Autoregressive Image Generation via Frequency Ordering
by: Huang, Zhihao, et al.
Published: (2025)
by: Huang, Zhihao, et al.
Published: (2025)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
by: Masrourisaadat, Nila, et al.
Published: (2024)
by: Masrourisaadat, Nila, et al.
Published: (2024)
Beyond Segmentation: Structurally Informed Facade Parsing from Imperfect Images
by: Janicki, Maciej, et al.
Published: (2026)
by: Janicki, Maciej, et al.
Published: (2026)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
by: Katakam, Raj Kiran Gupta
Published: (2026)
by: Katakam, Raj Kiran Gupta
Published: (2026)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
by: Zhu, Chenglin, et al.
Published: (2025)
by: Zhu, Chenglin, et al.
Published: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025)
by: Romero, Angel, et al.
Published: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
Similar Items
-
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
by: Yavari, Sara, et al.
Published: (2025) -
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
by: Zhou, Fang, et al.
Published: (2025) -
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
by: Eymaël, Alexandre, et al.
Published: (2024) -
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
by: Tang, Guowei
Published: (2026) -
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
by: Qian, Wenxu, et al.
Published: (2025)