Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zan, Zhang, Jingze, Chen, Yixin, Jia, Baoxiong, Liang, Wei, Huang, Siyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
by: Wang, Zan, et al.
Published: (2024)
by: Wang, Zan, et al.
Published: (2024)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
by: Lu, Ruijie, et al.
Published: (2024)
by: Lu, Ruijie, et al.
Published: (2024)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Scaling Up Dynamic Human-Scene Interaction Modeling
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
by: Linghu, Xiongkun, et al.
Published: (2026)
by: Linghu, Xiongkun, et al.
Published: (2026)
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
by: Zhang, Jinlu, et al.
Published: (2025)
by: Zhang, Jinlu, et al.
Published: (2025)
PhysPart: Physically Plausible Part Completion for Interactable Objects
by: Luo, Rundong, et al.
Published: (2024)
by: Luo, Rundong, et al.
Published: (2024)
Unifying 3D Vision-Language Understanding via Promptable Queries
by: Zhu, Ziyu, et al.
Published: (2024)
by: Zhu, Ziyu, et al.
Published: (2024)
DogMo: A Large-Scale Multi-View RGB-D Dataset for 4D Canine Motion Recovery
by: Wang, Zan, et al.
Published: (2025)
by: Wang, Zan, et al.
Published: (2025)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
by: Jia, Baoxiong, et al.
Published: (2024)
by: Jia, Baoxiong, et al.
Published: (2024)
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
by: Zhu, Ziyu, et al.
Published: (2025)
by: Zhu, Ziyu, et al.
Published: (2025)
SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
by: Yang, Yandan, et al.
Published: (2025)
by: Yang, Yandan, et al.
Published: (2025)
Dynamic Motion Blending for Versatile Motion Editing
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
by: Huang, Jiangyong, et al.
Published: (2025)
by: Huang, Jiangyong, et al.
Published: (2025)
Multi-modal Situated Reasoning in 3D Scenes
by: Linghu, Xiongkun, et al.
Published: (2024)
by: Linghu, Xiongkun, et al.
Published: (2024)
SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
by: Linghu, Xiongkun, et al.
Published: (2025)
by: Linghu, Xiongkun, et al.
Published: (2025)
GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
by: Lu, Guanxing, et al.
Published: (2025)
by: Lu, Guanxing, et al.
Published: (2025)
Task-oriented Sequential Grounding and Navigation in 3D Scenes
by: Zhang, Zhuofan, et al.
Published: (2024)
by: Zhang, Zhuofan, et al.
Published: (2024)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
by: Yang, Yandan, et al.
Published: (2024)
by: Yang, Yandan, et al.
Published: (2024)
Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
ARFlow: Human Action-Reaction Flow Matching with Physical Guidance
by: Jiang, Wentao, et al.
Published: (2025)
by: Jiang, Wentao, et al.
Published: (2025)
3D Scene Change Modeling With Consistent Multi-View Aggregation
by: Zhou, Zirui, et al.
Published: (2025)
by: Zhou, Zirui, et al.
Published: (2025)
Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis
by: Huang, Jiangyong, et al.
Published: (2025)
by: Huang, Jiangyong, et al.
Published: (2025)
Motion-Adaptive Multi-Scale Temporal Modelling with Skeleton-Constrained Spatial Graphs for Efficient 3D Human Pose Estimation
by: Li, Ruochen, et al.
Published: (2026)
by: Li, Ruochen, et al.
Published: (2026)
MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling
by: Yuan, Weihao, et al.
Published: (2024)
by: Yuan, Weihao, et al.
Published: (2024)
G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening
by: Yang, Zhiqi, et al.
Published: (2026)
by: Yang, Zhiqi, et al.
Published: (2026)
StyledStreets: Multi-style Street Simulator with Spatial and Temporal Consistency
by: Chen, Yuyin, et al.
Published: (2025)
by: Chen, Yuyin, et al.
Published: (2025)
Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal Interactions
by: Zheng, Yuanhong, et al.
Published: (2025)
by: Zheng, Yuanhong, et al.
Published: (2025)
One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
by: Yang, Siyuan, et al.
Published: (2023)
by: Yang, Siyuan, et al.
Published: (2023)
Human Motion Estimation with Everyday Wearables
by: Zhu, Siqi, et al.
Published: (2025)
by: Zhu, Siqi, et al.
Published: (2025)
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
by: Chen, Yixin, et al.
Published: (2026)
by: Chen, Yixin, et al.
Published: (2026)
Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
by: Wang, Zhumei, et al.
Published: (2025)
by: Wang, Zhumei, et al.
Published: (2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
HandX: Scaling Bimanual Motion and Interaction Generation
by: Zhang, Zimu, et al.
Published: (2026)
by: Zhang, Zimu, et al.
Published: (2026)
FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
R2G: Reasoning to Ground in 3D Scenes
by: Li, Yixuan, et al.
Published: (2024)
by: Li, Yixuan, et al.
Published: (2024)
ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
by: Zhao, Yanpeng, et al.
Published: (2026)
by: Zhao, Yanpeng, et al.
Published: (2026)
Spatial-Temporal State Propagation Autoregressive Model for 4D Object Generation
by: Yang, Liying, et al.
Published: (2026)
by: Yang, Liying, et al.
Published: (2026)
ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model
by: Chen, Wenshuo, et al.
Published: (2025)
by: Chen, Wenshuo, et al.
Published: (2025)
Similar Items
-
Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
by: Wang, Zan, et al.
Published: (2024) -
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
by: Liu, Yu, et al.
Published: (2024) -
MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
by: Lu, Ruijie, et al.
Published: (2024) -
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
by: Wang, Yan, et al.
Published: (2025) -
Scaling Up Dynamic Human-Scene Interaction Modeling
by: Jiang, Nan, et al.
Published: (2024)