POS: A Prompts Optimization Suite for Augmenting Text-to-Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Shijie, Xu, Huayi, Li, Mengjian, Geng, Weidong, Wang, Yaxiong, Wang, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Geometric-Photometric Joint Alignment for Facial Mesh Registration
von: Wang, Xizhi, et al.
Veröffentlicht: (2024)
von: Wang, Xizhi, et al.
Veröffentlicht: (2024)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
von: Ma, Yuting, et al.
Veröffentlicht: (2024)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
von: Huang, Ziqi, et al.
Veröffentlicht: (2024)
von: Huang, Ziqi, et al.
Veröffentlicht: (2024)
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
von: Wang, Qihang, et al.
Veröffentlicht: (2025)
Universal Prompt Optimizer for Safe Text-to-Image Generation
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
Optimizing Prompts for Text-to-Image Generation
von: Hao, Yaru, et al.
Veröffentlicht: (2022)
von: Hao, Yaru, et al.
Veröffentlicht: (2022)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
Consistency-aware Fake Videos Detection on Short Video Platforms
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
Text-Driven Diffusion Model for Sign Language Production
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
TIPO: Text to Image with Text Presampling for Prompt Optimization
von: Yeh, Shih-Ying, et al.
Veröffentlicht: (2024)
von: Yeh, Shih-Ying, et al.
Veröffentlicht: (2024)
Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization
von: Tan, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Tan, Xiaofeng, et al.
Veröffentlicht: (2024)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
VideoDirector: Precise Video Editing via Text-to-Video Models
von: Wang, Yukun, et al.
Veröffentlicht: (2024)
von: Wang, Yukun, et al.
Veröffentlicht: (2024)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
von: Wu, Shang, et al.
Veröffentlicht: (2026)
von: Wu, Shang, et al.
Veröffentlicht: (2026)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
von: Yang, Jiahui, et al.
Veröffentlicht: (2024)
von: Yang, Jiahui, et al.
Veröffentlicht: (2024)
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2024)
von: Kothandaraman, Divya, et al.
Veröffentlicht: (2024)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
von: Luo, Yang, et al.
Veröffentlicht: (2025)
von: Luo, Yang, et al.
Veröffentlicht: (2025)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
Motion Prompting: Controlling Video Generation with Motion Trajectories
von: Geng, Daniel, et al.
Veröffentlicht: (2024)
von: Geng, Daniel, et al.
Veröffentlicht: (2024)
TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
von: Huang, Yiyao, et al.
Veröffentlicht: (2025)
von: Huang, Yiyao, et al.
Veröffentlicht: (2025)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
von: Liu, Lingyu, et al.
Veröffentlicht: (2025)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
von: Wang, Jiarui, et al.
Veröffentlicht: (2025)
von: Wang, Jiarui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Geometric-Photometric Joint Alignment for Facial Mesh Registration
von: Wang, Xizhi, et al.
Veröffentlicht: (2024) -
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025) -
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2024) -
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
von: Yang, Shuyu, et al.
Veröffentlicht: (2025) -
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
von: Ma, Yuting, et al.
Veröffentlicht: (2024)