PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhan, Yu-Wei, Wang, Xin, Chen, Hong, Feng, Tongtong, Feng, Wei, Wang, Ren, Li, Guangyao, Li, Qing, Zhu, Wenwu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Multi-weather Cross-view Geo-localization Using Denoising Diffusion Models
par: Feng, Tongtong, et autres
Publié: (2024)
par: Feng, Tongtong, et autres
Publié: (2024)
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
par: Kong, Zhe, et autres
Publié: (2025)
par: Kong, Zhe, et autres
Publié: (2025)
Self-evolving Embodied AI
par: Feng, Tongtong, et autres
Publié: (2026)
par: Feng, Tongtong, et autres
Publié: (2026)
Multi-sentence Video Grounding for Long Video Generation
par: Feng, Wei, et autres
Publié: (2024)
par: Feng, Wei, et autres
Publié: (2024)
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
par: Li, Guangyao, et autres
Publié: (2026)
par: Li, Guangyao, et autres
Publié: (2026)
Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model
par: Zhou, Hang, et autres
Publié: (2024)
par: Zhou, Hang, et autres
Publié: (2024)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
par: Zhai, Yuanhao, et autres
Publié: (2024)
par: Zhai, Yuanhao, et autres
Publié: (2024)
EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks
par: Feng, Tongtong, et autres
Publié: (2025)
par: Feng, Tongtong, et autres
Publié: (2025)
BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
par: Zhan, Yu-Wei, et autres
Publié: (2025)
par: Zhan, Yu-Wei, et autres
Publié: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
par: Chen, Houlun, et autres
Publié: (2026)
par: Chen, Houlun, et autres
Publié: (2026)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
par: Wang, Jiankang, et autres
Publié: (2025)
par: Wang, Jiankang, et autres
Publié: (2025)
PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking
par: Bao, Jiacheng, et autres
Publié: (2026)
par: Bao, Jiacheng, et autres
Publié: (2026)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
par: Chen, Hong, et autres
Publié: (2024)
par: Chen, Hong, et autres
Publié: (2024)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
par: Chen, Houlun, et autres
Publié: (2024)
par: Chen, Houlun, et autres
Publié: (2024)
LLM4VG: Large Language Models Evaluation for Video Grounding
par: Feng, Wei, et autres
Publié: (2023)
par: Feng, Wei, et autres
Publié: (2023)
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
par: Wang, Boyuan, et autres
Publié: (2026)
par: Wang, Boyuan, et autres
Publié: (2026)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
par: Huang, Yidong, et autres
Publié: (2026)
par: Huang, Yidong, et autres
Publié: (2026)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
par: Gu, Jing, et autres
Publié: (2025)
par: Gu, Jing, et autres
Publié: (2025)
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
par: Liu, Huijie, et autres
Publié: (2025)
par: Liu, Huijie, et autres
Publié: (2025)
Disentangled Geometry and Appearance for Efficient Multi-View Surface Reconstruction and Rendering
par: Zhang, Qitong, et autres
Publié: (2025)
par: Zhang, Qitong, et autres
Publié: (2025)
4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
par: Lyu, Jin, et autres
Publié: (2026)
par: Lyu, Jin, et autres
Publié: (2026)
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
par: Su, Tongtong, et autres
Publié: (2025)
par: Su, Tongtong, et autres
Publié: (2025)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
par: Wang, Haoyu, et autres
Publié: (2025)
par: Wang, Haoyu, et autres
Publié: (2025)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
par: Xue, Qiyao, et autres
Publié: (2024)
par: Xue, Qiyao, et autres
Publié: (2024)
Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance Disentanglement
par: Yu, Runyi, et autres
Publié: (2024)
par: Yu, Runyi, et autres
Publié: (2024)
FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing
par: Yuan, Tianshuo, et autres
Publié: (2024)
par: Yuan, Tianshuo, et autres
Publié: (2024)
VideoPhy: Evaluating Physical Commonsense for Video Generation
par: Bansal, Hritik, et autres
Publié: (2024)
par: Bansal, Hritik, et autres
Publié: (2024)
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
par: Tang, Shengeng, et autres
Publié: (2024)
par: Tang, Shengeng, et autres
Publié: (2024)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
par: Cai, Yuanhao, et autres
Publié: (2025)
par: Cai, Yuanhao, et autres
Publié: (2025)
Toward Rich Video Human-Motion2D Generation
par: Xi, Ruihao, et autres
Publié: (2025)
par: Xi, Ruihao, et autres
Publié: (2025)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
par: Liu, Jinlin, et autres
Publié: (2024)
par: Liu, Jinlin, et autres
Publié: (2024)
ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
par: Zheng, Zhicheng, et autres
Publié: (2024)
par: Zheng, Zhicheng, et autres
Publié: (2024)
SpectralSplat: Appearance-Disentangled Feed-Forward Gaussian Splatting for Driving Scenes
par: Herau, Quentin, et autres
Publié: (2026)
par: Herau, Quentin, et autres
Publié: (2026)
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
par: Khalil, Ahmad, et autres
Publié: (2025)
par: Khalil, Ahmad, et autres
Publié: (2025)
MIMAFace: Face Animation via Motion-Identity Modulated Appearance Feature Learning
par: Han, Yue, et autres
Publié: (2024)
par: Han, Yue, et autres
Publié: (2024)
PhyRPR: Training-Free Physics-Constrained Video Generation
par: Zhao, Yibo, et autres
Publié: (2026)
par: Zhao, Yibo, et autres
Publié: (2026)
Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
par: Li, Xiao, et autres
Publié: (2025)
par: Li, Xiao, et autres
Publié: (2025)
PhiP-G: Physics-Guided Text-to-3D Compositional Scene Generation
par: Li, Qixuan, et autres
Publié: (2025)
par: Li, Qixuan, et autres
Publié: (2025)
Pose-Guided Fine-Grained Sign Language Video Generation
par: Shi, Tongkai, et autres
Publié: (2024)
par: Shi, Tongkai, et autres
Publié: (2024)
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
par: Gao, Mingju, et autres
Publié: (2026)
par: Gao, Mingju, et autres
Publié: (2026)
Documents similaires
-
Multi-weather Cross-view Geo-localization Using Denoising Diffusion Models
par: Feng, Tongtong, et autres
Publié: (2024) -
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
par: Kong, Zhe, et autres
Publié: (2025) -
Self-evolving Embodied AI
par: Feng, Tongtong, et autres
Publié: (2026) -
Multi-sentence Video Grounding for Long Video Generation
par: Feng, Wei, et autres
Publié: (2024) -
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
par: Li, Guangyao, et autres
Publié: (2026)