MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Pappa, Massimiliano, Collorone, Luca, Ficarra, Giovanni, Spinelli, Indro, Galasso, Fabio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
Length-Aware Motion Synthesis via Latent Diffusion
by: Sampieri, Alessio, et al.
Published: (2024)
by: Sampieri, Alessio, et al.
Published: (2024)
Social EgoMesh Estimation
by: Scofano, Luca, et al.
Published: (2024)
by: Scofano, Luca, et al.
Published: (2024)
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
by: Pappa, Massimiliano, et al.
Published: (2026)
by: Pappa, Massimiliano, et al.
Published: (2026)
Human Motion Unlearning
by: De Matteis, Edoardo, et al.
Published: (2025)
by: De Matteis, Edoardo, et al.
Published: (2025)
Following the Human Thread in Social Navigation
by: Scofano, Luca, et al.
Published: (2024)
by: Scofano, Luca, et al.
Published: (2024)
OVOSE: Open-Vocabulary Semantic Segmentation in Event-Based Cameras
by: Rahman, Muhammad Rameez Ur, et al.
Published: (2024)
by: Rahman, Muhammad Rameez Ur, et al.
Published: (2024)
Video Unlearning via Low-Rank Refusal Vector
by: Facchiano, Simone, et al.
Published: (2025)
by: Facchiano, Simone, et al.
Published: (2025)
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
by: Palma, Alessio, et al.
Published: (2026)
by: Palma, Alessio, et al.
Published: (2026)
Quantifying Self-Preservation Bias in Large Language Models
by: Migliarini, Matteo, et al.
Published: (2026)
by: Migliarini, Matteo, et al.
Published: (2026)
Adaptive Point Transformer
by: Baiocchi, Alessandro, et al.
Published: (2024)
by: Baiocchi, Alessandro, et al.
Published: (2024)
Fine-grained text-driven dual-human motion generation via dynamic hierarchical interaction
by: Li, Mu, et al.
Published: (2025)
by: Li, Mu, et al.
Published: (2025)
ViPO: Visual Preference Optimization at Scale
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
Altogether: Image Captioning via Re-aligning Alt-text
by: Xu, Hu, et al.
Published: (2024)
by: Xu, Hu, et al.
Published: (2024)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization
by: Yang, Yixuan, et al.
Published: (2025)
by: Yang, Yixuan, et al.
Published: (2025)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
SeRpEnt: Selective Resampling for Expressive State Space Models
by: Rando, Stefano, et al.
Published: (2025)
by: Rando, Stefano, et al.
Published: (2025)
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
by: Sarkar, Pritam, et al.
Published: (2025)
by: Sarkar, Pritam, et al.
Published: (2025)
About latent roles in forecasting players in team sports
by: Scofano, Luca, et al.
Published: (2023)
by: Scofano, Luca, et al.
Published: (2023)
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
by: Zhang, Yangsong, et al.
Published: (2026)
by: Zhang, Yangsong, et al.
Published: (2026)
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
by: Tang, Jiyang, et al.
Published: (2025)
by: Tang, Jiyang, et al.
Published: (2025)
Text-driven 3D Human Generation via Contrastive Preference Optimization
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
Rethinking Direct Preference Optimization in Diffusion Models
by: Kang, Junyong, et al.
Published: (2025)
by: Kang, Junyong, et al.
Published: (2025)
Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
by: Lee, Ji Soo, et al.
Published: (2025)
by: Lee, Ji Soo, et al.
Published: (2025)
InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment
by: Lu, Yunhong, et al.
Published: (2025)
by: Lu, Yunhong, et al.
Published: (2025)
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
by: Liu, Chaohu, et al.
Published: (2025)
by: Liu, Chaohu, et al.
Published: (2025)
$\text{Di}^2\text{Pose}$: Discrete Diffusion Model for Occluded 3D Human Pose Estimation
by: Wang, Weiquan, et al.
Published: (2024)
by: Wang, Weiquan, et al.
Published: (2024)
SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions
by: Sbrolli, Cristian, et al.
Published: (2025)
by: Sbrolli, Cristian, et al.
Published: (2025)
Discriminator-Free Direct Preference Optimization for Video Diffusion
by: Cheng, Haoran, et al.
Published: (2025)
by: Cheng, Haoran, et al.
Published: (2025)
ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization
by: Shen, Wenhao, et al.
Published: (2025)
by: Shen, Wenhao, et al.
Published: (2025)
Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
by: Cui, Jiahao, et al.
Published: (2025)
by: Cui, Jiahao, et al.
Published: (2025)
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
by: Kim, Dohyun, et al.
Published: (2025)
by: Kim, Dohyun, et al.
Published: (2025)
NewMove: Customizing text-to-video models with novel motions
by: Materzynska, Joanna, et al.
Published: (2023)
by: Materzynska, Joanna, et al.
Published: (2023)
An extremely coarse feedback signal is sufficient for learning human-aligned visual representations
by: Mehta, Yash, et al.
Published: (2026)
by: Mehta, Yash, et al.
Published: (2026)
Di3PO - Diptych Diffusion DPO for Targeted Improvements in Image Generation
by: Reddy, Sanjana, et al.
Published: (2026)
by: Reddy, Sanjana, et al.
Published: (2026)
GenMask: Adapting DiT for Segmentation via Direct Mask Generation
by: Yang, Yuhuan, et al.
Published: (2026)
by: Yang, Yuhuan, et al.
Published: (2026)
Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization
by: Su, Weijian, et al.
Published: (2026)
by: Su, Weijian, et al.
Published: (2026)
Similar Items
-
MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
by: Collorone, Luca, et al.
Published: (2025) -
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
by: Collorone, Luca, et al.
Published: (2025) -
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation
by: Collorone, Luca, et al.
Published: (2025) -
Length-Aware Motion Synthesis via Latent Diffusion
by: Sampieri, Alessio, et al.
Published: (2024) -
Social EgoMesh Estimation
by: Scofano, Luca, et al.
Published: (2024)