VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Jiale, Lyu, Ruiliang, Gu, Xiaotao, Liu, Xiao, Xu, Jiazheng, Lu, Yida, Teng, Jiayan, Yang, Zhuoyi, Dong, Yuxiao, Tang, Jie, Wang, Hongning, Huang, Minlie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
von: Cheng, Jiale, et al.
Veröffentlicht: (2023)
von: Cheng, Jiale, et al.
Veröffentlicht: (2023)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)
von: Lu, Yida, et al.
Veröffentlicht: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
von: Liu, Shizhan, et al.
Veröffentlicht: (2025)
von: Liu, Shizhan, et al.
Veröffentlicht: (2025)
Glyph: Scaling Context Windows via Visual-Text Compression
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
von: Zheng, Wendi, et al.
Veröffentlicht: (2024)
von: Zheng, Wendi, et al.
Veröffentlicht: (2024)
Concat-ID: Towards Universal Identity-Preserving Video Synthesis
von: Zhong, Yong, et al.
Veröffentlicht: (2025)
von: Zhong, Yong, et al.
Veröffentlicht: (2025)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2025)
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
von: Xu, Jiazheng, et al.
Veröffentlicht: (2024)
von: Xu, Jiazheng, et al.
Veröffentlicht: (2024)
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
von: Hong, Wenyi, et al.
Veröffentlicht: (2025)
von: Hong, Wenyi, et al.
Veröffentlicht: (2025)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
von: Yan, Wenhao, et al.
Veröffentlicht: (2025)
von: Yan, Wenhao, et al.
Veröffentlicht: (2025)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
von: Ke, Pei, et al.
Veröffentlicht: (2023)
von: Ke, Pei, et al.
Veröffentlicht: (2023)
Language Model Decoding as Direct Metrics Optimization
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
Towards Efficient Exact Optimization of Language Model Alignment
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
VPO: Leveraging the Number of Votes in Preference Optimization
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
von: Huang, Zhenpeng, et al.
Veröffentlicht: (2026)
von: Huang, Zhenpeng, et al.
Veröffentlicht: (2026)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Data Selection via Optimal Control for Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
LVBench: An Extreme Long Video Understanding Benchmark
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
Benchmarking Complex Instruction-Following with Multiple Constraints Composition
von: Wen, Bosi, et al.
Veröffentlicht: (2024)
von: Wen, Bosi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
von: Cheng, Jiale, et al.
Veröffentlicht: (2023) -
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024) -
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024) -
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025) -
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)