LibraGen: Playing a Balance Game in Subject-Driven Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Jiahao, Lao, Shanshan, Liu, Lijie, Li, Gen, Qi, Tianhao, Han, Wei, Li, Bingchuan, Liu, Fangfang, Chen, Zhuowei, Ma, Tianxiang, HE, Qian, Zhou, Yi, Xie, Xiaohua
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912970579116032
author Zhu, Jiahao
Lao, Shanshan
Liu, Lijie
Li, Gen
Qi, Tianhao
Han, Wei
Li, Bingchuan
Liu, Fangfang
Chen, Zhuowei
Ma, Tianxiang
HE, Qian
Zhou, Yi
Xie, Xiaohua
author_facet Zhu, Jiahao
Lao, Shanshan
Liu, Lijie
Li, Gen
Qi, Tianhao
Han, Wei
Li, Bingchuan
Liu, Fangfang
Chen, Zhuowei
Ma, Tianxiang
HE, Qian
Zhou, Yi
Xie, Xiaohua
contents With the advancement of video generation foundation models (VGFMs), customized generation, particularly subject-to-video (S2V), has attracted growing attention. However, a key challenge lies in balancing the intrinsic priors of a VGFM, such as motion coherence, visual aesthetics, and prompt alignment, with its newly derived S2V capability. Existing methods often neglect this balance by enhancing one aspect at the expense of others. To address this, we propose LibraGen, a novel framework that views extending foundation models for S2V generation as a balance game between intrinsic VGFM strengths and S2V capability. Specifically, guided by the core philosophy of "Raising the Fulcrum, Tuning to Balance," we identify data quality as the fulcrum and advocate a quality-over-quantity approach. We construct a hybrid pipeline that combines automated and manual data filtering to improve overall data quality. To further harmonize the VGFM's native capabilities with its S2V extension, we introduce a Tune-to-Balance post-training paradigm. During supervised fine-tuning, both cross-pair and in-pair data are incorporated, and model merging is employed to achieve an effective trade-off. Subsequently, two tailored direct preference optimization (DPO) pipelines, namely Consis-DPO and Real-Fake DPO, are designed and merged to consolidate this balance. During inference, we introduce a time-dependent dynamic classifier-free guidance scheme to enable flexible and fine-grained control. Experimental results demonstrate that LibraGen outperforms both open-source and commercial S2V models using only thousand-scale training data.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LibraGen: Playing a Balance Game in Subject-Driven Video Generation
Zhu, Jiahao
Lao, Shanshan
Liu, Lijie
Li, Gen
Qi, Tianhao
Han, Wei
Li, Bingchuan
Liu, Fangfang
Chen, Zhuowei
Ma, Tianxiang
HE, Qian
Zhou, Yi
Xie, Xiaohua
Computer Vision and Pattern Recognition
With the advancement of video generation foundation models (VGFMs), customized generation, particularly subject-to-video (S2V), has attracted growing attention. However, a key challenge lies in balancing the intrinsic priors of a VGFM, such as motion coherence, visual aesthetics, and prompt alignment, with its newly derived S2V capability. Existing methods often neglect this balance by enhancing one aspect at the expense of others. To address this, we propose LibraGen, a novel framework that views extending foundation models for S2V generation as a balance game between intrinsic VGFM strengths and S2V capability. Specifically, guided by the core philosophy of "Raising the Fulcrum, Tuning to Balance," we identify data quality as the fulcrum and advocate a quality-over-quantity approach. We construct a hybrid pipeline that combines automated and manual data filtering to improve overall data quality. To further harmonize the VGFM's native capabilities with its S2V extension, we introduce a Tune-to-Balance post-training paradigm. During supervised fine-tuning, both cross-pair and in-pair data are incorporated, and model merging is employed to achieve an effective trade-off. Subsequently, two tailored direct preference optimization (DPO) pipelines, namely Consis-DPO and Real-Fake DPO, are designed and merged to consolidate this balance. During inference, we introduce a time-dependent dynamic classifier-free guidance scheme to enable flexible and fine-grained control. Experimental results demonstrate that LibraGen outperforms both open-source and commercial S2V models using only thousand-scale training data.
title LibraGen: Playing a Balance Game in Subject-Driven Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.13506