LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zekun, An, Sizhe, Tang, Chengcheng, Guo, Chuan, Shugurov, Ivan, Zhang, Linguang, Zhao, Amy, Sridhar, Srinath, Tao, Lingling, Mittal, Abhay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PHD: Personalized 3D Human Body Fitting with Point Diffusion
by: Ho, Hsuan-I, et al.
Published: (2025)
by: Ho, Hsuan-I, et al.
Published: (2025)
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
by: Cong, Xiaoyan, et al.
Published: (2026)
by: Cong, Xiaoyan, et al.
Published: (2026)
LLaMo: Large Language Model-based Molecular Graph Assistant
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
IAM: Identity-Aware Human Motion and Shape Joint Generation
by: Jia, Wenqi, et al.
Published: (2026)
by: Jia, Wenqi, et al.
Published: (2026)
FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
by: Chen, Kefan, et al.
Published: (2024)
by: Chen, Kefan, et al.
Published: (2024)
MotionGlot: A Multi-Embodied Motion Generation Model
by: Harithas, Sudarshan, et al.
Published: (2024)
by: Harithas, Sudarshan, et al.
Published: (2024)
TokenUnify: Scaling Up Autoregressive Pretraining for Neuron Segmentation
by: Chen, Yinda, et al.
Published: (2024)
by: Chen, Yinda, et al.
Published: (2024)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation
by: Hwang, Inwoo, et al.
Published: (2026)
by: Hwang, Inwoo, et al.
Published: (2026)
Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
by: Ma, Chenglong, et al.
Published: (2025)
by: Ma, Chenglong, et al.
Published: (2025)
Art3D: Training-Free 3D Generation from Flat-Colored Illustration
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
GenHeld: Generating and Editing Handheld Objects
by: Min, Chaerin, et al.
Published: (2024)
by: Min, Chaerin, et al.
Published: (2024)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
by: Rai, Aashish, et al.
Published: (2024)
by: Rai, Aashish, et al.
Published: (2024)
HuMoCon: Concept Discovery for Human Motion Understanding
by: Fang, Qihang, et al.
Published: (2025)
by: Fang, Qihang, et al.
Published: (2025)
ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
by: Lu, Shunlin, et al.
Published: (2024)
by: Lu, Shunlin, et al.
Published: (2024)
GenHSI: Controllable Generation of Human-Scene Interaction Videos
by: Li, Zekun, et al.
Published: (2025)
by: Li, Zekun, et al.
Published: (2025)
Geometric Neural Distance Fields for Learning Human Motion Priors
by: Yu, Zhengdi, et al.
Published: (2025)
by: Yu, Zhengdi, et al.
Published: (2025)
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2024)
by: Fan, Lijie, et al.
Published: (2024)
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
by: Braun, Tobias, et al.
Published: (2026)
by: Braun, Tobias, et al.
Published: (2026)
UniMo: Unified Motion Generation and Understanding with Chain of Thought
by: Wang, Guocun, et al.
Published: (2026)
by: Wang, Guocun, et al.
Published: (2026)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
by: Wang, Hongpeng, et al.
Published: (2026)
by: Wang, Hongpeng, et al.
Published: (2026)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
by: Rai, Aashish, et al.
Published: (2026)
by: Rai, Aashish, et al.
Published: (2026)
CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes
by: Xu, Xianghao, et al.
Published: (2024)
by: Xu, Xianghao, et al.
Published: (2024)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
by: NextStep Team, et al.
Published: (2025)
by: NextStep Team, et al.
Published: (2025)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
MoSa: Motion Generation with Scalable Autoregressive Modeling
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Unified Cross-Scale 3D Generation and Understanding via Autoregressive Modeling
by: Lu, Shuqi, et al.
Published: (2025)
by: Lu, Shuqi, et al.
Published: (2025)
Spatial Audio Motion Understanding and Reasoning
by: Sridhar, Arvind Krishna, et al.
Published: (2025)
by: Sridhar, Arvind Krishna, et al.
Published: (2025)
Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
by: Tang, Yutao, et al.
Published: (2025)
by: Tang, Yutao, et al.
Published: (2025)
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
by: Kanjula, Karthik Reddy, et al.
Published: (2025)
by: Kanjula, Karthik Reddy, et al.
Published: (2025)
MANUS: Markerless Grasp Capture using Articulated 3D Gaussians
by: Pokhariya, Chandradeep, et al.
Published: (2023)
by: Pokhariya, Chandradeep, et al.
Published: (2023)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
by: Ke, Guolin, et al.
Published: (2025)
by: Ke, Guolin, et al.
Published: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Fractal Autoregressive Depth Estimation with Continuous Token Diffusion
by: Zhang, Jinchang, et al.
Published: (2026)
by: Zhang, Jinchang, et al.
Published: (2026)
Similar Items
-
PHD: Personalized 3D Human Body Fitting with Point Diffusion
by: Ho, Hsuan-I, et al.
Published: (2025) -
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
by: Cong, Xiaoyan, et al.
Published: (2026) -
LLaMo: Large Language Model-based Molecular Graph Assistant
by: Park, Jinyoung, et al.
Published: (2024) -
IAM: Identity-Aware Human Motion and Shape Joint Generation
by: Jia, Wenqi, et al.
Published: (2026) -
FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
by: Chen, Kefan, et al.
Published: (2024)