EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Ruibing, Zhou, Mingyue, Gui, Yuwei, Luo, Mingshuang, Ma, Bingpeng, Chang, Hong, Shan, Shiguang, Chen, Xilin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025)
by: Hou, Ruibing, et al.
Published: (2025)
Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation
by: Liang, Jiachen, et al.
Published: (2024)
by: Liang, Jiachen, et al.
Published: (2024)
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
by: Luo, Mingshuang, et al.
Published: (2024)
by: Luo, Mingshuang, et al.
Published: (2024)
Clothes-Changing Person Re-Identification with Feasibility-Aware Intermediary Matching
by: Zhao, Jiahe, et al.
Published: (2024)
by: Zhao, Jiahe, et al.
Published: (2024)
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
by: Li, Yinqi, et al.
Published: (2025)
by: Li, Yinqi, et al.
Published: (2025)
UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models
by: Liang, Jiachen, et al.
Published: (2024)
by: Liang, Jiachen, et al.
Published: (2024)
MATS: An Audio Language Model under Text-only Supervision
by: Wang, Wen, et al.
Published: (2025)
by: Wang, Wen, et al.
Published: (2025)
KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
by: Yang, Zaifei, et al.
Published: (2025)
by: Yang, Zaifei, et al.
Published: (2025)
Morph: A Motion-free Physics Optimization Framework for Human Motion Generation
by: Li, Zhuo, et al.
Published: (2024)
by: Li, Zhuo, et al.
Published: (2024)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
Component-Based Out-of-Distribution Detection
by: Liu, Wenrui, et al.
Published: (2026)
by: Liu, Wenrui, et al.
Published: (2026)
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
by: Luo, Mingshuang, et al.
Published: (2026)
by: Luo, Mingshuang, et al.
Published: (2026)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
by: Li, Yinqi, et al.
Published: (2025)
by: Li, Yinqi, et al.
Published: (2025)
Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
by: Liang, Jiachen, et al.
Published: (2025)
by: Liang, Jiachen, et al.
Published: (2025)
EgoLM: Multi-Modal Language Model of Egocentric Motions
by: Hong, Fangzhou, et al.
Published: (2024)
by: Hong, Fangzhou, et al.
Published: (2024)
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models
by: Cai, Yufei, et al.
Published: (2025)
by: Cai, Yufei, et al.
Published: (2025)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios
by: Huang, Jie, et al.
Published: (2024)
by: Huang, Jie, et al.
Published: (2024)
UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
by: Li, Keliang, et al.
Published: (2024)
by: Li, Keliang, et al.
Published: (2024)
EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing
by: Hwang, Inwoo, et al.
Published: (2026)
by: Hwang, Inwoo, et al.
Published: (2026)
Task Attribute Distance for Few-Shot Learning: Theoretical Analysis and Applications
by: Hu, Minyang, et al.
Published: (2024)
by: Hu, Minyang, et al.
Published: (2024)
VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task
by: Long, Xingming, et al.
Published: (2025)
by: Long, Xingming, et al.
Published: (2025)
Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
by: Cho, Kyungwon, et al.
Published: (2025)
by: Cho, Kyungwon, et al.
Published: (2025)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers
by: Yuan, Zheng, et al.
Published: (2024)
by: Yuan, Zheng, et al.
Published: (2024)
EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation
by: Saroha, Abhishek, et al.
Published: (2026)
by: Saroha, Abhishek, et al.
Published: (2026)
Multi-PA: A Multi-perspective Benchmark on Privacy Assessment for Large Vision-Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2024)
by: Wang, Zhongqi, et al.
Published: (2024)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary Diffusion
by: Liu, Mengdi, et al.
Published: (2025)
by: Liu, Mengdi, et al.
Published: (2025)
Empathetic Motion Generation for Humanoid Educational Robots via Reasoning-Guided Vision--Language--Motion Diffusion Architecture
by: Sun, Fuze, et al.
Published: (2026)
by: Sun, Fuze, et al.
Published: (2026)
Learning Separable Hidden Unit Contributions for Speaker-Adaptive Lip-Reading
by: Luo, Songtao, et al.
Published: (2023)
by: Luo, Songtao, et al.
Published: (2023)
Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR
by: Li, Zhenyu, et al.
Published: (2026)
by: Li, Zhenyu, et al.
Published: (2026)
EgoHDM: An Online Egocentric-Inertial Human Motion Capture, Localization, and Dense Mapping System
by: Liu, Bonan, et al.
Published: (2024)
by: Liu, Bonan, et al.
Published: (2024)
Similar Items
-
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025) -
Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation
by: Liang, Jiachen, et al.
Published: (2024) -
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
by: Luo, Mingshuang, et al.
Published: (2024) -
Clothes-Changing Person Re-Identification with Feasibility-Aware Intermediary Matching
by: Zhao, Jiahe, et al.
Published: (2024) -
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
by: Li, Yinqi, et al.
Published: (2025)