AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Haonan, Wu, Xiangzuo, Liao, Huan, Xu, Zunnan, Hu, Zhongyuan, Li, Ronghui, Zhang, Yachao, Li, Xiu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AToM: Amortized Text-to-Mesh using 2D Diffusion
di: Qian, Guocheng, et al.
Pubblicazione: (2024)
di: Qian, Guocheng, et al.
Pubblicazione: (2024)
AToM: Adaptive Theory-of-Mind-Based Human Motion Prediction in Long-Term Human-Robot Interactions
di: Liao, Yuwen, et al.
Pubblicazione: (2025)
di: Liao, Yuwen, et al.
Pubblicazione: (2025)
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
di: Xu, Zunnan, et al.
Pubblicazione: (2024)
di: Xu, Zunnan, et al.
Pubblicazione: (2024)
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
di: Ding, Wei, et al.
Pubblicazione: (2024)
di: Ding, Wei, et al.
Pubblicazione: (2024)
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
di: Liu, Zihao, et al.
Pubblicazione: (2025)
di: Liu, Zihao, et al.
Pubblicazione: (2025)
BATON: Aligning Text-to-Audio Model with Human Preference Feedback
di: Liao, Huan, et al.
Pubblicazione: (2024)
di: Liao, Huan, et al.
Pubblicazione: (2024)
Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
di: Lin, Yukang, et al.
Pubblicazione: (2023)
di: Lin, Yukang, et al.
Pubblicazione: (2023)
Conversión de diagramas de procesos en diagramas de casos de usos usando AToM3
di: CARLOS ALBERTO ÁLVAREZ
Pubblicazione: (2005)
di: CARLOS ALBERTO ÁLVAREZ
Pubblicazione: (2005)
Activity-based and agent-based Transport model of Melbourne (AToM): an open multi-modal transport simulation model for Greater Melbourne
di: Jafari, Afshin, et al.
Pubblicazione: (2021)
di: Jafari, Afshin, et al.
Pubblicazione: (2021)
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
di: Zhang, Youliang, et al.
Pubblicazione: (2024)
di: Zhang, Youliang, et al.
Pubblicazione: (2024)
Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
di: Gong, Chaoqun, et al.
Pubblicazione: (2024)
di: Gong, Chaoqun, et al.
Pubblicazione: (2024)
REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment
di: Han, Haonan, et al.
Pubblicazione: (2024)
di: Han, Haonan, et al.
Pubblicazione: (2024)
MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
di: Wu, Xiangzuo, et al.
Pubblicazione: (2025)
di: Wu, Xiangzuo, et al.
Pubblicazione: (2025)
Exploring Multi-Modal Control in Music-Driven Dance Generation
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning
di: Liu, Xiaoyang, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyang, et al.
Pubblicazione: (2024)
Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models
di: Ye, Zhankai, et al.
Pubblicazione: (2026)
di: Ye, Zhankai, et al.
Pubblicazione: (2026)
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
di: Li, Ronghui, et al.
Pubblicazione: (2026)
di: Li, Ronghui, et al.
Pubblicazione: (2026)
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
Adversarial Training on Purification (AToP): Advancing Both Robustness and Generalization
di: Lin, Guang, et al.
Pubblicazione: (2024)
di: Lin, Guang, et al.
Pubblicazione: (2024)
ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
di: Weng, Wanjiang, et al.
Pubblicazione: (2025)
di: Weng, Wanjiang, et al.
Pubblicazione: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
di: Li, Xiaojie, et al.
Pubblicazione: (2025)
di: Li, Xiaojie, et al.
Pubblicazione: (2025)
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
di: Weng, Wanjiang, et al.
Pubblicazione: (2025)
di: Weng, Wanjiang, et al.
Pubblicazione: (2025)
WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
di: Ying, Qijun, et al.
Pubblicazione: (2025)
di: Ying, Qijun, et al.
Pubblicazione: (2025)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
di: Chen, Junying, et al.
Pubblicazione: (2025)
di: Chen, Junying, et al.
Pubblicazione: (2025)
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
di: Lv, Jiaxi, et al.
Pubblicazione: (2023)
di: Lv, Jiaxi, et al.
Pubblicazione: (2023)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
di: Cui, Mingxuan, et al.
Pubblicazione: (2026)
di: Cui, Mingxuan, et al.
Pubblicazione: (2026)
Insertion site profiling analysis method for three active R2 retrotransposons in human cells
di: Wu, Yachao
Pubblicazione: (2026)
di: Wu, Yachao
Pubblicazione: (2026)
Event-T2M: Event-level Conditioning for Complex Text-to-Motion Synthesis
di: Hong, Seong-Eun, et al.
Pubblicazione: (2026)
di: Hong, Seong-Eun, et al.
Pubblicazione: (2026)
MGStream: Motion-aware 3D Gaussian for Streamable Dynamic Scene Reconstruction
di: Bao, Zhenyu, et al.
Pubblicazione: (2025)
di: Bao, Zhenyu, et al.
Pubblicazione: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2023)
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2023)
Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
di: Lu, Qiang, et al.
Pubblicazione: (2025)
di: Lu, Qiang, et al.
Pubblicazione: (2025)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
di: Jiang, Songtao, et al.
Pubblicazione: (2025)
di: Jiang, Songtao, et al.
Pubblicazione: (2025)
TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
di: Mao, Dongxing, et al.
Pubblicazione: (2026)
di: Mao, Dongxing, et al.
Pubblicazione: (2026)
PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
di: Li, Shufan, et al.
Pubblicazione: (2024)
di: Li, Shufan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AToM: Amortized Text-to-Mesh using 2D Diffusion
di: Qian, Guocheng, et al.
Pubblicazione: (2024) -
AToM: Adaptive Theory-of-Mind-Based Human Motion Prediction in Long-Term Human-Robot Interactions
di: Liao, Yuwen, et al.
Pubblicazione: (2025) -
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
di: Xu, Zunnan, et al.
Pubblicazione: (2024) -
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
di: Ding, Wei, et al.
Pubblicazione: (2024) -
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
di: Liu, Zihao, et al.
Pubblicazione: (2025)