AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
Fuente:
arXiv
Guardado en:
| Autores principales: | Han, Haonan, Wu, Xiangzuo, Liao, Huan, Xu, Zunnan, Hu, Zhongyuan, Li, Ronghui, Zhang, Yachao, Li, Xiu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AToM: Amortized Text-to-Mesh using 2D Diffusion
por: Qian, Guocheng, et al.
Publicado: (2024)
por: Qian, Guocheng, et al.
Publicado: (2024)
AToM: Adaptive Theory-of-Mind-Based Human Motion Prediction in Long-Term Human-Robot Interactions
por: Liao, Yuwen, et al.
Publicado: (2025)
por: Liao, Yuwen, et al.
Publicado: (2025)
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
por: Xu, Zunnan, et al.
Publicado: (2024)
por: Xu, Zunnan, et al.
Publicado: (2024)
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
por: Ding, Wei, et al.
Publicado: (2024)
por: Ding, Wei, et al.
Publicado: (2024)
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
por: Liu, Zihao, et al.
Publicado: (2025)
por: Liu, Zihao, et al.
Publicado: (2025)
BATON: Aligning Text-to-Audio Model with Human Preference Feedback
por: Liao, Huan, et al.
Publicado: (2024)
por: Liao, Huan, et al.
Publicado: (2024)
Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
por: Lin, Yukang, et al.
Publicado: (2023)
por: Lin, Yukang, et al.
Publicado: (2023)
Conversión de diagramas de procesos en diagramas de casos de usos usando AToM3
por: CARLOS ALBERTO ÁLVAREZ
Publicado: (2005)
por: CARLOS ALBERTO ÁLVAREZ
Publicado: (2005)
Activity-based and agent-based Transport model of Melbourne (AToM): an open multi-modal transport simulation model for Greater Melbourne
por: Jafari, Afshin, et al.
Publicado: (2021)
por: Jafari, Afshin, et al.
Publicado: (2021)
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
por: Zhang, Youliang, et al.
Publicado: (2024)
por: Zhang, Youliang, et al.
Publicado: (2024)
Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
por: Gong, Chaoqun, et al.
Publicado: (2024)
por: Gong, Chaoqun, et al.
Publicado: (2024)
REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment
por: Han, Haonan, et al.
Publicado: (2024)
por: Han, Haonan, et al.
Publicado: (2024)
MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
por: Wu, Xiangzuo, et al.
Publicado: (2025)
por: Wu, Xiangzuo, et al.
Publicado: (2025)
Exploring Multi-Modal Control in Music-Driven Dance Generation
por: Li, Ronghui, et al.
Publicado: (2024)
por: Li, Ronghui, et al.
Publicado: (2024)
MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning
por: Liu, Xiaoyang, et al.
Publicado: (2024)
por: Liu, Xiaoyang, et al.
Publicado: (2024)
Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation
por: Huang, Jiaqi, et al.
Publicado: (2025)
por: Huang, Jiaqi, et al.
Publicado: (2025)
GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models
por: Ye, Zhankai, et al.
Publicado: (2026)
por: Ye, Zhankai, et al.
Publicado: (2026)
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
por: Li, Ronghui, et al.
Publicado: (2026)
por: Li, Ronghui, et al.
Publicado: (2026)
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
por: Li, Ronghui, et al.
Publicado: (2024)
por: Li, Ronghui, et al.
Publicado: (2024)
Adversarial Training on Purification (AToP): Advancing Both Robustness and Generalization
por: Lin, Guang, et al.
Publicado: (2024)
por: Lin, Guang, et al.
Publicado: (2024)
ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
por: Weng, Wanjiang, et al.
Publicado: (2025)
por: Weng, Wanjiang, et al.
Publicado: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
por: Wu, Tong, et al.
Publicado: (2024)
por: Wu, Tong, et al.
Publicado: (2024)
Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
por: Li, Xiaojie, et al.
Publicado: (2025)
por: Li, Xiaojie, et al.
Publicado: (2025)
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
por: Li, Ronghui, et al.
Publicado: (2024)
por: Li, Ronghui, et al.
Publicado: (2024)
InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions
por: Li, Ronghui, et al.
Publicado: (2024)
por: Li, Ronghui, et al.
Publicado: (2024)
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
por: Wang, Jiaxuan, et al.
Publicado: (2026)
por: Wang, Jiaxuan, et al.
Publicado: (2026)
SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
por: Huang, Jiaqi, et al.
Publicado: (2025)
por: Huang, Jiaqi, et al.
Publicado: (2025)
ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
por: Weng, Wanjiang, et al.
Publicado: (2025)
por: Weng, Wanjiang, et al.
Publicado: (2025)
WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
por: Ying, Qijun, et al.
Publicado: (2025)
por: Ying, Qijun, et al.
Publicado: (2025)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
por: Chen, Junying, et al.
Publicado: (2025)
por: Chen, Junying, et al.
Publicado: (2025)
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
por: Lv, Jiaxi, et al.
Publicado: (2023)
por: Lv, Jiaxi, et al.
Publicado: (2023)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
por: Cui, Mingxuan, et al.
Publicado: (2026)
por: Cui, Mingxuan, et al.
Publicado: (2026)
Insertion site profiling analysis method for three active R2 retrotransposons in human cells
por: Wu, Yachao
Publicado: (2026)
por: Wu, Yachao
Publicado: (2026)
Event-T2M: Event-level Conditioning for Complex Text-to-Motion Synthesis
por: Hong, Seong-Eun, et al.
Publicado: (2026)
por: Hong, Seong-Eun, et al.
Publicado: (2026)
MGStream: Motion-aware 3D Gaussian for Streamable Dynamic Scene Reconstruction
por: Bao, Zhenyu, et al.
Publicado: (2025)
por: Bao, Zhenyu, et al.
Publicado: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
por: Prabhudesai, Mihir, et al.
Publicado: (2023)
por: Prabhudesai, Mihir, et al.
Publicado: (2023)
Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
por: Lu, Qiang, et al.
Publicado: (2025)
por: Lu, Qiang, et al.
Publicado: (2025)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
por: Jiang, Songtao, et al.
Publicado: (2025)
por: Jiang, Songtao, et al.
Publicado: (2025)
TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
por: Mao, Dongxing, et al.
Publicado: (2026)
por: Mao, Dongxing, et al.
Publicado: (2026)
PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
por: Li, Shufan, et al.
Publicado: (2024)
por: Li, Shufan, et al.
Publicado: (2024)
Ejemplares similares
-
AToM: Amortized Text-to-Mesh using 2D Diffusion
por: Qian, Guocheng, et al.
Publicado: (2024) -
AToM: Adaptive Theory-of-Mind-Based Human Motion Prediction in Long-Term Human-Robot Interactions
por: Liao, Yuwen, et al.
Publicado: (2025) -
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
por: Xu, Zunnan, et al.
Publicado: (2024) -
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
por: Ding, Wei, et al.
Publicado: (2024) -
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
por: Liu, Zihao, et al.
Publicado: (2025)