Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rahman, Zillur, Sheng, Alex, Meo, Cristian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
von: Wu, Shang, et al.
Veröffentlicht: (2026)
von: Wu, Shang, et al.
Veröffentlicht: (2026)
Video-T1: Test-Time Scaling for Video Generation
von: Liu, Fangfu, et al.
Veröffentlicht: (2025)
von: Liu, Fangfu, et al.
Veröffentlicht: (2025)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
von: Wang, Ni, et al.
Veröffentlicht: (2024)
von: Wang, Ni, et al.
Veröffentlicht: (2024)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
von: Jeong, Suchae, et al.
Veröffentlicht: (2025)
von: Jeong, Suchae, et al.
Veröffentlicht: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
von: Mo, Wenyi, et al.
Veröffentlicht: (2024)
von: Mo, Wenyi, et al.
Veröffentlicht: (2024)
Scaling Image and Video Generation via Test-Time Evolutionary Search
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation
von: Ye, Wen, et al.
Veröffentlicht: (2025)
von: Ye, Wen, et al.
Veröffentlicht: (2025)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
von: Rahman, Kazi Mahathir, et al.
Veröffentlicht: (2025)
von: Rahman, Kazi Mahathir, et al.
Veröffentlicht: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
von: Kim, Subin, et al.
Veröffentlicht: (2025)
von: Kim, Subin, et al.
Veröffentlicht: (2025)
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
Minority-Focused Text-to-Image Generation via Prompt Optimization
von: Um, Soobin, et al.
Veröffentlicht: (2024)
von: Um, Soobin, et al.
Veröffentlicht: (2024)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
von: Huang, Jen-Yuan, et al.
Veröffentlicht: (2026)
von: Huang, Jen-Yuan, et al.
Veröffentlicht: (2026)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
von: Erregue, Iñaki, et al.
Veröffentlicht: (2026)
von: Erregue, Iñaki, et al.
Veröffentlicht: (2026)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
von: Xue, Qiyao, et al.
Veröffentlicht: (2024)
von: Xue, Qiyao, et al.
Veröffentlicht: (2024)
Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval
von: Yang, Junkai, et al.
Veröffentlicht: (2026)
von: Yang, Junkai, et al.
Veröffentlicht: (2026)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
von: Cho, CH, et al.
Veröffentlicht: (2025)
von: Cho, CH, et al.
Veröffentlicht: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
Video-As-Prompt: Unified Semantic Control for Video Generation
von: Bian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Bian, Yuxuan, et al.
Veröffentlicht: (2025)
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
von: Xu, Hang, et al.
Veröffentlicht: (2025)
von: Xu, Hang, et al.
Veröffentlicht: (2025)
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
Singular Value Scaling: Efficient Generative Model Compression via Pruned Weights Refinement
von: Kim, Hyeonjin, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonjin, et al.
Veröffentlicht: (2024)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
von: Du, Yang, et al.
Veröffentlicht: (2025)
von: Du, Yang, et al.
Veröffentlicht: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
von: Kang, Peng, et al.
Veröffentlicht: (2025)
von: Kang, Peng, et al.
Veröffentlicht: (2025)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2026)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
DOTA: Deformable Optimized Transformer Architecture for End-to-End Text Recognition with Retrieval-Augmented Generation
von: Nithisopa, Naphat, et al.
Veröffentlicht: (2025)
von: Nithisopa, Naphat, et al.
Veröffentlicht: (2025)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Batch-Instructed Gradient for Prompt Evolution:Systematic Prompt Optimization for Enhanced Text-to-Image Synthesis
von: Yang, Xinrui, et al.
Veröffentlicht: (2024)
von: Yang, Xinrui, et al.
Veröffentlicht: (2024)
GenDDS: Generating Diverse Driving Video Scenarios with Prompt-to-Video Generative Model
von: Fu, Yongjie, et al.
Veröffentlicht: (2024)
von: Fu, Yongjie, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation
von: Molino, Daniele, et al.
Veröffentlicht: (2026)
von: Molino, Daniele, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
von: Wu, Shang, et al.
Veröffentlicht: (2026) -
Video-T1: Test-Time Scaling for Video Generation
von: Liu, Fangfu, et al.
Veröffentlicht: (2025) -
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
von: Paul, Dhiman, et al.
Veröffentlicht: (2024) -
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
von: Wang, Ni, et al.
Veröffentlicht: (2024) -
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
von: Jeong, Suchae, et al.
Veröffentlicht: (2025)