Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Furuta, Hiroki, Zen, Heiga, Schuurmans, Dale, Faust, Aleksandra, Matsuo, Yutaka, Liang, Percy, Yang, Sherry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
Geometric-Averaged Preference Optimization for Soft Preference Labels
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
by: Furuta, Hiroki, et al.
Published: (2023)
by: Furuta, Hiroki, et al.
Published: (2023)
Video as the New Language for Real-World Decision Making
by: Yang, Sherry, et al.
Published: (2024)
by: Yang, Sherry, et al.
Published: (2024)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
by: Oshima, Yuta, et al.
Published: (2024)
by: Oshima, Yuta, et al.
Published: (2024)
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
by: Furuta, Ryosuke, et al.
Published: (2023)
by: Furuta, Ryosuke, et al.
Published: (2023)
Improving Video Generation with Human Feedback
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
by: Choi, Minkyu, et al.
Published: (2025)
by: Choi, Minkyu, et al.
Published: (2025)
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
by: Han, Guangyi, et al.
Published: (2025)
by: Han, Guangyi, et al.
Published: (2025)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
by: Cai, Xinhao, et al.
Published: (2025)
by: Cai, Xinhao, et al.
Published: (2025)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
by: Gur, Izzeddin, et al.
Published: (2023)
by: Gur, Izzeddin, et al.
Published: (2023)
ClawCraneNet: Leveraging Object-level Relation for Text-based Video Segmentation
by: Liang, Chen, et al.
Published: (2021)
by: Liang, Chen, et al.
Published: (2021)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
by: Sakai, Yuki, et al.
Published: (2025)
by: Sakai, Yuki, et al.
Published: (2025)
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
by: Zhang, Jinlu, et al.
Published: (2025)
by: Zhang, Jinlu, et al.
Published: (2025)
Rich Human Feedback for Text-to-Image Generation
by: Liang, Youwei, et al.
Published: (2023)
by: Liang, Youwei, et al.
Published: (2023)
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
by: Guan, Jiazhi, et al.
Published: (2026)
by: Guan, Jiazhi, et al.
Published: (2026)
Multimodal Priors-Augmented Text-Driven 3D Human-Object Interaction Generation
by: Wang, Yin, et al.
Published: (2026)
by: Wang, Yin, et al.
Published: (2026)
AGFSync: Leveraging AI-Generated Feedback for Preference Optimization in Text-to-Image Generation
by: An, Jingkun, et al.
Published: (2024)
by: An, Jingkun, et al.
Published: (2024)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
by: Xu, Sirui, et al.
Published: (2024)
by: Xu, Sirui, et al.
Published: (2024)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
by: Taniguchi, Takara, et al.
Published: (2024)
by: Taniguchi, Takara, et al.
Published: (2024)
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
MOVi: Training-free Text-conditioned Multi-Object Video Generation
by: Rahman, Aimon, et al.
Published: (2025)
by: Rahman, Aimon, et al.
Published: (2025)
EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
by: Cha, Junuk, et al.
Published: (2024)
by: Cha, Junuk, et al.
Published: (2024)
InTraGen: Trajectory-controlled Video Generation for Object Interactions
by: Liu, Zuhao, et al.
Published: (2024)
by: Liu, Zuhao, et al.
Published: (2024)
Aligning Anime Video Generation with Human Feedback
by: Zhu, Bingwen, et al.
Published: (2025)
by: Zhu, Bingwen, et al.
Published: (2025)
Improving Unsupervised Video Object Segmentation via Fake Flow Generation
by: Cho, Suhwan, et al.
Published: (2024)
by: Cho, Suhwan, et al.
Published: (2024)
Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
by: Lan, Bangxiang, et al.
Published: (2025)
by: Lan, Bangxiang, et al.
Published: (2025)
Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation
by: Liang, Chen, et al.
Published: (2021)
by: Liang, Chen, et al.
Published: (2021)
InterFusion: Text-Driven Generation of 3D Human-Object Interaction
by: Dai, Sisi, et al.
Published: (2024)
by: Dai, Sisi, et al.
Published: (2024)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
by: Oh, Yoonjin, et al.
Published: (2025)
by: Oh, Yoonjin, et al.
Published: (2025)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
VideoTetris: Towards Compositional Text-to-Video Generation
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
by: Yang, Shiyuan, et al.
Published: (2024)
by: Yang, Shiyuan, et al.
Published: (2024)
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
by: Liao, Mingxiang, et al.
Published: (2024)
by: Liao, Mingxiang, et al.
Published: (2024)
Similar Items
-
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
by: Oshima, Yuta, et al.
Published: (2025) -
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
by: Oshima, Yuta, et al.
Published: (2025) -
Geometric-Averaged Preference Optimization for Soft Preference Labels
by: Furuta, Hiroki, et al.
Published: (2024) -
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
by: Oshima, Yuta, et al.
Published: (2025) -
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
by: Furuta, Hiroki, et al.
Published: (2023)