AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Haonan, Wu, Xiangzuo, Liao, Huan, Xu, Zunnan, Hu, Zhongyuan, Li, Ronghui, Zhang, Yachao, Li, Xiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AToM: Amortized Text-to-Mesh using 2D Diffusion
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
AToM: Adaptive Theory-of-Mind-Based Human Motion Prediction in Long-Term Human-Robot Interactions
von: Liao, Yuwen, et al.
Veröffentlicht: (2025)
von: Liao, Yuwen, et al.
Veröffentlicht: (2025)
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
von: Xu, Zunnan, et al.
Veröffentlicht: (2024)
von: Xu, Zunnan, et al.
Veröffentlicht: (2024)
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
von: Ding, Wei, et al.
Veröffentlicht: (2024)
von: Ding, Wei, et al.
Veröffentlicht: (2024)
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
BATON: Aligning Text-to-Audio Model with Human Preference Feedback
von: Liao, Huan, et al.
Veröffentlicht: (2024)
von: Liao, Huan, et al.
Veröffentlicht: (2024)
Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
von: Lin, Yukang, et al.
Veröffentlicht: (2023)
von: Lin, Yukang, et al.
Veröffentlicht: (2023)
Conversión de diagramas de procesos en diagramas de casos de usos usando AToM3
von: CARLOS ALBERTO ÁLVAREZ
Veröffentlicht: (2005)
von: CARLOS ALBERTO ÁLVAREZ
Veröffentlicht: (2005)
Activity-based and agent-based Transport model of Melbourne (AToM): an open multi-modal transport simulation model for Greater Melbourne
von: Jafari, Afshin, et al.
Veröffentlicht: (2021)
von: Jafari, Afshin, et al.
Veröffentlicht: (2021)
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
von: Zhang, Youliang, et al.
Veröffentlicht: (2024)
von: Zhang, Youliang, et al.
Veröffentlicht: (2024)
Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
von: Gong, Chaoqun, et al.
Veröffentlicht: (2024)
von: Gong, Chaoqun, et al.
Veröffentlicht: (2024)
REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment
von: Han, Haonan, et al.
Veröffentlicht: (2024)
von: Han, Haonan, et al.
Veröffentlicht: (2024)
MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
von: Wu, Xiangzuo, et al.
Veröffentlicht: (2025)
von: Wu, Xiangzuo, et al.
Veröffentlicht: (2025)
Exploring Multi-Modal Control in Music-Driven Dance Generation
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation
von: Huang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Huang, Jiaqi, et al.
Veröffentlicht: (2025)
GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models
von: Ye, Zhankai, et al.
Veröffentlicht: (2026)
von: Ye, Zhankai, et al.
Veröffentlicht: (2026)
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
von: Li, Ronghui, et al.
Veröffentlicht: (2026)
von: Li, Ronghui, et al.
Veröffentlicht: (2026)
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
Adversarial Training on Purification (AToP): Advancing Both Robustness and Generalization
von: Lin, Guang, et al.
Veröffentlicht: (2024)
von: Lin, Guang, et al.
Veröffentlicht: (2024)
ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
von: Weng, Wanjiang, et al.
Veröffentlicht: (2025)
von: Weng, Wanjiang, et al.
Veröffentlicht: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
von: Wang, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Wang, Jiaxuan, et al.
Veröffentlicht: (2026)
SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
von: Huang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Huang, Jiaqi, et al.
Veröffentlicht: (2025)
ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
von: Weng, Wanjiang, et al.
Veröffentlicht: (2025)
von: Weng, Wanjiang, et al.
Veröffentlicht: (2025)
WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
von: Ying, Qijun, et al.
Veröffentlicht: (2025)
von: Ying, Qijun, et al.
Veröffentlicht: (2025)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
von: Chen, Junying, et al.
Veröffentlicht: (2025)
von: Chen, Junying, et al.
Veröffentlicht: (2025)
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
von: Lv, Jiaxi, et al.
Veröffentlicht: (2023)
von: Lv, Jiaxi, et al.
Veröffentlicht: (2023)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
von: Cui, Mingxuan, et al.
Veröffentlicht: (2026)
von: Cui, Mingxuan, et al.
Veröffentlicht: (2026)
Insertion site profiling analysis method for three active R2 retrotransposons in human cells
von: Wu, Yachao
Veröffentlicht: (2026)
von: Wu, Yachao
Veröffentlicht: (2026)
Event-T2M: Event-level Conditioning for Complex Text-to-Motion Synthesis
von: Hong, Seong-Eun, et al.
Veröffentlicht: (2026)
von: Hong, Seong-Eun, et al.
Veröffentlicht: (2026)
MGStream: Motion-aware 3D Gaussian for Streamable Dynamic Scene Reconstruction
von: Bao, Zhenyu, et al.
Veröffentlicht: (2025)
von: Bao, Zhenyu, et al.
Veröffentlicht: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
von: Lu, Qiang, et al.
Veröffentlicht: (2025)
von: Lu, Qiang, et al.
Veröffentlicht: (2025)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
von: Mao, Dongxing, et al.
Veröffentlicht: (2026)
von: Mao, Dongxing, et al.
Veröffentlicht: (2026)
PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
von: Li, Shufan, et al.
Veröffentlicht: (2024)
von: Li, Shufan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AToM: Amortized Text-to-Mesh using 2D Diffusion
von: Qian, Guocheng, et al.
Veröffentlicht: (2024) -
AToM: Adaptive Theory-of-Mind-Based Human Motion Prediction in Long-Term Human-Robot Interactions
von: Liao, Yuwen, et al.
Veröffentlicht: (2025) -
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
von: Xu, Zunnan, et al.
Veröffentlicht: (2024) -
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
von: Ding, Wei, et al.
Veröffentlicht: (2024) -
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
von: Liu, Zihao, et al.
Veröffentlicht: (2025)