Imbalance in Balance: Online Concept Balancing in Generation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Yukai, Ou, Jiarong, Chen, Rui, Yang, Haotian, Wang, Jiahao, Tao, Xin, Wan, Pengfei, Zhang, Di, Gai, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
by: Wang, Qiuheng, et al.
Published: (2024)
by: Wang, Qiuheng, et al.
Published: (2024)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
by: Wang, Ruotong, et al.
Published: (2025)
by: Wang, Ruotong, et al.
Published: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
by: Hou, Liang, et al.
Published: (2025)
by: Hou, Liang, et al.
Published: (2025)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
by: Huang, Yuzhou, et al.
Published: (2025)
by: Huang, Yuzhou, et al.
Published: (2025)
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
by: Li, Ouxiang, et al.
Published: (2025)
by: Li, Ouxiang, et al.
Published: (2025)
Towards Precise Scaling Laws for Video Diffusion Transformers
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection
by: Ding, Kaixin, et al.
Published: (2025)
by: Ding, Kaixin, et al.
Published: (2025)
Balanced Sharpness-Aware Minimization for Imbalanced Regression
by: Liu, Yahao, et al.
Published: (2025)
by: Liu, Yahao, et al.
Published: (2025)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
by: Chen, Junyi, et al.
Published: (2026)
by: Chen, Junyi, et al.
Published: (2026)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
by: Cheng, Junhao, et al.
Published: (2026)
by: Cheng, Junhao, et al.
Published: (2026)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
CamCloneMaster: Enabling Reference-based Camera Control for Video Generation
by: Luo, Yawen, et al.
Published: (2025)
by: Luo, Yawen, et al.
Published: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
BaSAL: Size-Balanced Warm Start Active Learning for LiDAR Semantic Segmentation
by: Wei, Jiarong, et al.
Published: (2023)
by: Wei, Jiarong, et al.
Published: (2023)
Bringing Balance to Hand Shape Classification: Mitigating Data Imbalance Through Generative Models
by: Rios, Gaston Gustavo, et al.
Published: (2025)
by: Rios, Gaston Gustavo, et al.
Published: (2025)
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
by: Yang, Zhenhao, et al.
Published: (2026)
by: Yang, Zhenhao, et al.
Published: (2026)
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation
by: Wang, Qinghe, et al.
Published: (2025)
by: Wang, Qinghe, et al.
Published: (2025)
A Survey of Interactive Generative Video
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Rates
by: Nguyen, Phuong-Anh, et al.
Published: (2026)
by: Nguyen, Phuong-Anh, et al.
Published: (2026)
UniVideo: Unified Understanding, Generation, and Editing for Videos
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Generalization of Diffusion Models Arises with a Balanced Representation Space
by: Zhang, Zekai, et al.
Published: (2025)
by: Zhang, Zekai, et al.
Published: (2025)
VideoTetris: Towards Compositional Text-to-Video Generation
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
VRMM: A Volumetric Relightable Morphable Head Model
by: Yang, Haotian, et al.
Published: (2024)
by: Yang, Haotian, et al.
Published: (2024)
Scaling Image and Video Generation via Test-Time Evolutionary Search
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
by: Xu, Yiyan, et al.
Published: (2026)
by: Xu, Yiyan, et al.
Published: (2026)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
by: Fang, Zhixue, et al.
Published: (2026)
by: Fang, Zhixue, et al.
Published: (2026)
Balanced 3DGS: Gaussian-wise Parallelism Rendering with Fine-Grained Tiling
by: Gui, Hao, et al.
Published: (2024)
by: Gui, Hao, et al.
Published: (2024)
Alleviating Class Imbalance in Semi-supervised Multi-organ Segmentation via Balanced Subclass Regularization
by: Feng, Zhenghao, et al.
Published: (2024)
by: Feng, Zhenghao, et al.
Published: (2024)
BD-KD: Balancing the Divergences for Online Knowledge Distillation
by: Amara, Ibtihel, et al.
Published: (2022)
by: Amara, Ibtihel, et al.
Published: (2022)
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
by: Wang, Qinghe, et al.
Published: (2025)
by: Wang, Qinghe, et al.
Published: (2025)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
by: Zhu, Jiahao, et al.
Published: (2026)
by: Zhu, Jiahao, et al.
Published: (2026)
NeRF Inpainting with Geometric Diffusion Prior and Balanced Score Distillation
by: Zhang, Menglin, et al.
Published: (2024)
by: Zhang, Menglin, et al.
Published: (2024)
Addressing Imbalanced Domain-Incremental Learning through Dual-Balance Collaborative Experts
by: Li, Lan, et al.
Published: (2025)
by: Li, Lan, et al.
Published: (2025)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
NegVSR: Augmenting Negatives for Generalized Noise Modeling in Real-World Video Super-Resolution
by: Song, Yexing, et al.
Published: (2023)
by: Song, Yexing, et al.
Published: (2023)
Similar Items
-
Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
by: Wang, Qiuheng, et al.
Published: (2024) -
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
by: Wang, Ruotong, et al.
Published: (2025) -
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
by: Hou, Liang, et al.
Published: (2025) -
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
by: Huang, Yuzhou, et al.
Published: (2025) -
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
by: Li, Ouxiang, et al.
Published: (2025)