The Unified Balance Theory of Second-Moment Exponential Scaling Optimizers in Visual Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Gongyue, Liu, Honghai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Natural Spectral Fusion: p-Exponent Cyclic Scheduling and Early Decision-Boundary Alignment in First-Order Optimization
by: Zhang, Gongyue, et al.
Published: (2025)
by: Zhang, Gongyue, et al.
Published: (2025)
FUSE: First-Order and Second-Order Unified SynthEsis in Stochastic Optimization
by: Jiang, Zhanhong, et al.
Published: (2025)
by: Jiang, Zhanhong, et al.
Published: (2025)
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
by: Bai, Hao, et al.
Published: (2026)
by: Bai, Hao, et al.
Published: (2026)
Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream
by: Gokce, Abdulkadir, et al.
Published: (2024)
by: Gokce, Abdulkadir, et al.
Published: (2024)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
by: Hoffmann, David T., et al.
Published: (2023)
by: Hoffmann, David T., et al.
Published: (2023)
Many-Task Federated Fine-Tuning via Unified Task Vectors
by: Tsouvalas, Vasileios, et al.
Published: (2025)
by: Tsouvalas, Vasileios, et al.
Published: (2025)
Break a Lag: Triple Exponential Moving Average for Enhanced Optimization
by: Peleg, Roi, et al.
Published: (2023)
by: Peleg, Roi, et al.
Published: (2023)
Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation
by: Sun, Jun, et al.
Published: (2025)
by: Sun, Jun, et al.
Published: (2025)
Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
by: Guan, Yunchuan, et al.
Published: (2025)
by: Guan, Yunchuan, et al.
Published: (2025)
Understanding Bias in Large-Scale Visual Datasets
by: Zeng, Boya, et al.
Published: (2024)
by: Zeng, Boya, et al.
Published: (2024)
Quantifying Task Priority for Multi-Task Optimization
by: Jeong, Wooseong, et al.
Published: (2024)
by: Jeong, Wooseong, et al.
Published: (2024)
Hierarchical Invariance for Robust and Interpretable Vision Tasks at Larger Scales
by: Qi, Shuren, et al.
Published: (2024)
by: Qi, Shuren, et al.
Published: (2024)
Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
db-SP: Accelerating Sparse Attention for Visual Generative Models with Dual-Balanced Sequence Parallelism
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
Adaptive Parametric Activation: Unifying and Generalising Activation Functions Across Tasks
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
Balancing the Scales: Enhancing Fairness in Facial Expression Recognition with Latent Alignment
by: Rizvi, Syed Sameen Ahmad, et al.
Published: (2024)
by: Rizvi, Syed Sameen Ahmad, et al.
Published: (2024)
Federated Learning with Uncertainty and Personalization via Efficient Second-order Optimization
by: Pal, Shivam, et al.
Published: (2024)
by: Pal, Shivam, et al.
Published: (2024)
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
by: Wu, Yecheng, et al.
Published: (2024)
by: Wu, Yecheng, et al.
Published: (2024)
SGW-based Multi-Task Learning in Vision Tasks
by: Zhang, Ruiyuan, et al.
Published: (2024)
by: Zhang, Ruiyuan, et al.
Published: (2024)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
by: Safwan, Itbaan, et al.
Published: (2025)
by: Safwan, Itbaan, et al.
Published: (2025)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
by: Wang, Xiyao, et al.
Published: (2025)
by: Wang, Xiyao, et al.
Published: (2025)
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
by: Akl, Ahmed, et al.
Published: (2024)
by: Akl, Ahmed, et al.
Published: (2024)
Glyph: Scaling Context Windows via Visual-Text Compression
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
Direct Discriminative Optimization: Your Likelihood-Based Visual Generative Model is Secretly a GAN Discriminator
by: Zheng, Kaiwen, et al.
Published: (2025)
by: Zheng, Kaiwen, et al.
Published: (2025)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)
by: Neuhaus, Yannic, et al.
Published: (2026)
GO4Align: Group Optimization for Multi-Task Alignment
by: Shen, Jiayi, et al.
Published: (2024)
by: Shen, Jiayi, et al.
Published: (2024)
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
Downstream Task Guided Masking Learning in Masked Autoencoders Using Multi-Level Optimization
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts
by: Lou, Meng, et al.
Published: (2026)
by: Lou, Meng, et al.
Published: (2026)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Gradient-based Class Weighting for Unsupervised Domain Adaptation in Dense Prediction Visual Tasks
by: Alcover-Couso, Roberto, et al.
Published: (2024)
by: Alcover-Couso, Roberto, et al.
Published: (2024)
Revisiting Disentanglement in Downstream Tasks: A Study on Its Necessity for Abstract Visual Reasoning
by: Nai, Ruiqian, et al.
Published: (2024)
by: Nai, Ruiqian, et al.
Published: (2024)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
by: Li, Bingxin
Published: (2025)
by: Li, Bingxin
Published: (2025)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Unified Classification and Rejection: A One-versus-All Framework
by: Cheng, Zhen, et al.
Published: (2023)
by: Cheng, Zhen, et al.
Published: (2023)
Adaptive Moments are Surprisingly Effective for Plug-and-Play Diffusion Sampling
by: Belardi, Christian, et al.
Published: (2026)
by: Belardi, Christian, et al.
Published: (2026)
Similar Items
-
Natural Spectral Fusion: p-Exponent Cyclic Scheduling and Early Decision-Boundary Alignment in First-Order Optimization
by: Zhang, Gongyue, et al.
Published: (2025) -
FUSE: First-Order and Second-Order Unified SynthEsis in Stochastic Optimization
by: Jiang, Zhanhong, et al.
Published: (2025) -
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
by: Bai, Hao, et al.
Published: (2026) -
Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream
by: Gokce, Abdulkadir, et al.
Published: (2024) -
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
by: Hoffmann, David T., et al.
Published: (2023)