A Reason-then-Describe Instruction Interpreter for Controllable Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shengqiong, Ye, Weicai, Zhang, Yuanxing, Wang, Jiahao, Liu, Quande, Wang, Xintao, Wan, Pengfei, Gai, Kun, Fei, Hao, Chua, Tat-Seng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908675867672576
author Wu, Shengqiong
Ye, Weicai
Zhang, Yuanxing
Wang, Jiahao
Liu, Quande
Wang, Xintao
Wan, Pengfei
Gai, Kun
Fei, Hao
Chua, Tat-Seng
author_facet Wu, Shengqiong
Ye, Weicai
Zhang, Yuanxing
Wang, Jiahao
Liu, Quande
Wang, Xintao
Wan, Pengfei
Gai, Kun
Fei, Hao
Chua, Tat-Seng
contents Diffusion Transformers have significantly improved video fidelity and temporal coherence, however, practical controllability remains limited. Concise, ambiguous, and compositionally complex user inputs contrast with the detailed prompts used in training, yielding an intent-output mismatch. We propose ReaDe, a universal, model-agnostic interpreter that converts raw instructions into precise, actionable specifications for downstream video generators. ReaDe follows a reason-then-describe paradigm: it first analyzes the user request to identify core requirements and resolve ambiguities, then produces detailed guidance that enables faithful, controllable generation. We train ReaDe via a two-stage optimization: (i) reasoning-augmented supervision imparts analytic parsing with stepwise traces and dense captions, and (ii) a multi-dimensional reward assigner enables stable, feedback-driven refinement for natural-style captions. Experiments across single- and multi-condition scenarios show consistent gains in instruction fidelity, caption accuracy, and downstream video quality, with strong generalization to reasoning-intensive and unseen inputs. ReaDe offers a practical route to aligning controllable video generation with accurately interpreted user intent. Project Page: https://sqwu.top/ReaDe/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20563
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
Wu, Shengqiong
Ye, Weicai
Zhang, Yuanxing
Wang, Jiahao
Liu, Quande
Wang, Xintao
Wan, Pengfei
Gai, Kun
Fei, Hao
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Diffusion Transformers have significantly improved video fidelity and temporal coherence, however, practical controllability remains limited. Concise, ambiguous, and compositionally complex user inputs contrast with the detailed prompts used in training, yielding an intent-output mismatch. We propose ReaDe, a universal, model-agnostic interpreter that converts raw instructions into precise, actionable specifications for downstream video generators. ReaDe follows a reason-then-describe paradigm: it first analyzes the user request to identify core requirements and resolve ambiguities, then produces detailed guidance that enables faithful, controllable generation. We train ReaDe via a two-stage optimization: (i) reasoning-augmented supervision imparts analytic parsing with stepwise traces and dense captions, and (ii) a multi-dimensional reward assigner enables stable, feedback-driven refinement for natural-style captions. Experiments across single- and multi-condition scenarios show consistent gains in instruction fidelity, caption accuracy, and downstream video quality, with strong generalization to reasoning-intensive and unseen inputs. ReaDe offers a practical route to aligning controllable video generation with accurately interpreted user intent. Project Page: https://sqwu.top/ReaDe/.
title A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20563