OpenT2M: No-frill Motion Generation with Open-source,Large-scale, High-quality Data
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Bin, Zheng, Sipeng, Luo, Hao, Li, Boyuan, Liu, Jing, Lu, Zongqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Motion Generation using Part-level Reliable Data from Videos
by: Li, Boyuan, et al.
Published: (2025)
by: Li, Boyuan, et al.
Published: (2025)
Scaling Large Motion Models with Million-Level Human Motions
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
by: Cao, Bin, et al.
Published: (2025)
by: Cao, Bin, et al.
Published: (2025)
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
by: Liu, Jiazheng, et al.
Published: (2025)
by: Liu, Jiazheng, et al.
Published: (2025)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)
by: Zheng, Sipeng, et al.
Published: (2024)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
by: Zhou, Bohan, et al.
Published: (2025)
by: Zhou, Bohan, et al.
Published: (2025)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
by: Gupta, Tanmay, et al.
Published: (2026)
by: Gupta, Tanmay, et al.
Published: (2026)
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
by: Luo, Hao, et al.
Published: (2024)
by: Luo, Hao, et al.
Published: (2024)
OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
by: Zhang, Jinlu, et al.
Published: (2025)
by: Zhang, Jinlu, et al.
Published: (2025)
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
Programmable Motion Generation for Open-Set Motion Control Tasks
by: Liu, Hanchao, et al.
Published: (2024)
by: Liu, Hanchao, et al.
Published: (2024)
Open-sourced Data Ecosystem in Autonomous Driving: the Present and Future
by: Li, Hongyang, et al.
Published: (2023)
by: Li, Hongyang, et al.
Published: (2023)
Advancing Open-source World Models
by: Robbyant Team, et al.
Published: (2026)
by: Robbyant Team, et al.
Published: (2026)
XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
Visual Grounding for Object-Level Generalization in Reinforcement Learning
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
OpenMaterial: A Large-scale Dataset of Complex Materials for 3D Reconstruction
by: Dang, Zheng, et al.
Published: (2024)
by: Dang, Zheng, et al.
Published: (2024)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement
by: Zheng, Jiakun, et al.
Published: (2026)
by: Zheng, Jiakun, et al.
Published: (2026)
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
by: Nan, Kepan, et al.
Published: (2024)
by: Nan, Kepan, et al.
Published: (2024)
Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
by: Lin, Jing, et al.
Published: (2023)
by: Lin, Jing, et al.
Published: (2023)
RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation
by: Cao, Boyuan, et al.
Published: (2024)
by: Cao, Boyuan, et al.
Published: (2024)
OpenMEDLab: An Open-source Platform for Multi-modality Foundation Models in Medicine
by: Wang, Xiaosong, et al.
Published: (2024)
by: Wang, Xiaosong, et al.
Published: (2024)
OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers
by: Liang, Han, et al.
Published: (2023)
by: Liang, Han, et al.
Published: (2023)
DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
by: Ye, Zilyu, et al.
Published: (2024)
by: Ye, Zilyu, et al.
Published: (2024)
Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data
by: Fan, Ke, et al.
Published: (2025)
by: Fan, Ke, et al.
Published: (2025)
OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map Construction
by: Zhao, Hongbo, et al.
Published: (2024)
by: Zhao, Hongbo, et al.
Published: (2024)
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
by: Chi, Haohan, et al.
Published: (2025)
by: Chi, Haohan, et al.
Published: (2025)
GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts
by: Milacski, Zoltán Á., et al.
Published: (2024)
by: Milacski, Zoltán Á., et al.
Published: (2024)
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
DisMo: Disentangled Motion Representations for Open-World Motion Transfer
by: Ressler-Antal, Thomas, et al.
Published: (2025)
by: Ressler-Antal, Thomas, et al.
Published: (2025)
Towards Open Domain Text-Driven Synthesis of Multi-Person Motions
by: Shan, Mengyi, et al.
Published: (2024)
by: Shan, Mengyi, et al.
Published: (2024)
OpenUrban3D: Annotation-Free Open-Vocabulary Semantic Segmentation of Large-Scale Urban Point Clouds
by: Wang, Chongyu, et al.
Published: (2025)
by: Wang, Chongyu, et al.
Published: (2025)
Wan: Open and Advanced Large-Scale Video Generative Models
by: Wan, Team, et al.
Published: (2025)
by: Wan, Team, et al.
Published: (2025)
Similar Items
-
Robust Motion Generation using Part-level Reliable Data from Videos
by: Li, Boyuan, et al.
Published: (2025) -
Scaling Large Motion Models with Million-Level Human Motions
by: Wang, Ye, et al.
Published: (2024) -
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
by: Cao, Bin, et al.
Published: (2025) -
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
by: Liu, Jiazheng, et al.
Published: (2025) -
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)