ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Meiqi, Zhu, Jiashu, Feng, Xiaokun, Chen, Chubin, Zhu, Chen, Song, Bingze, Mao, Fangyuan, Wu, Jiahong, Chu, Xiangxiang, Huang, Kaiqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915569941348352
author Wu, Meiqi
Zhu, Jiashu
Feng, Xiaokun
Chen, Chubin
Zhu, Chen
Song, Bingze
Mao, Fangyuan
Wu, Jiahong
Chu, Xiangxiang
Huang, Kaiqi
author_facet Wu, Meiqi
Zhu, Jiashu
Feng, Xiaokun
Chen, Chubin
Zhu, Chen
Song, Bingze
Mao, Fangyuan
Wu, Jiahong
Chu, Xiangxiang
Huang, Kaiqi
contents Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a prompt-guided adaptive test-time search strategy that dynamically adjusts both the inference search space and reward function according to semantic relationships in the prompt. This enables more coherent and visually plausible videos in challenging imaginative settings. To evaluate progress in this direction, we introduce LDT-Bench, the first dedicated benchmark for long-distance semantic prompts, consisting of 2,839 diverse concept pairs and an automated protocol for assessing creative generation capabilities. Extensive experiments show that ImagerySearch consistently outperforms strong video generation baselines and existing test-time scaling approaches on LDT-Bench, and achieves competitive improvements on VBench, demonstrating its effectiveness across diverse prompt types. We will release LDT-Bench and code to facilitate future research on imaginative video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
Wu, Meiqi
Zhu, Jiashu
Feng, Xiaokun
Chen, Chubin
Zhu, Chen
Song, Bingze
Mao, Fangyuan
Wu, Jiahong
Chu, Xiangxiang
Huang, Kaiqi
Computer Vision and Pattern Recognition
Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a prompt-guided adaptive test-time search strategy that dynamically adjusts both the inference search space and reward function according to semantic relationships in the prompt. This enables more coherent and visually plausible videos in challenging imaginative settings. To evaluate progress in this direction, we introduce LDT-Bench, the first dedicated benchmark for long-distance semantic prompts, consisting of 2,839 diverse concept pairs and an automated protocol for assessing creative generation capabilities. Extensive experiments show that ImagerySearch consistently outperforms strong video generation baselines and existing test-time scaling approaches on LDT-Bench, and achieves competitive improvements on VBench, demonstrating its effectiveness across diverse prompt types. We will release LDT-Bench and code to facilitate future research on imaginative video generation.
title ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14847