InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yangsong, Butt, Abdul Ahad, Varol, Gül, Laptev, Ivan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
di: Zhang, Yangsong, et al.
Pubblicazione: (2026)
di: Zhang, Yangsong, et al.
Pubblicazione: (2026)
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
di: Xu, Sirui, et al.
Pubblicazione: (2025)
di: Xu, Sirui, et al.
Pubblicazione: (2025)
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
di: Bensabath, Léore, et al.
Pubblicazione: (2024)
di: Bensabath, Léore, et al.
Pubblicazione: (2024)
SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation
di: Athanasiou, Nikos, et al.
Pubblicazione: (2023)
di: Athanasiou, Nikos, et al.
Pubblicazione: (2023)
Learning text-to-video retrieval from image captioning
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
di: Xu, Sirui, et al.
Pubblicazione: (2026)
di: Xu, Sirui, et al.
Pubblicazione: (2026)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
di: Souček, Tomáš, et al.
Pubblicazione: (2023)
di: Souček, Tomáš, et al.
Pubblicazione: (2023)
InterFusion: Text-Driven Generation of 3D Human-Object Interaction
di: Dai, Sisi, et al.
Pubblicazione: (2024)
di: Dai, Sisi, et al.
Pubblicazione: (2024)
InterTrack: Tracking Human Object Interaction without Object Templates
di: Xie, Xianghui, et al.
Pubblicazione: (2024)
di: Xie, Xianghui, et al.
Pubblicazione: (2024)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
di: Romero, David, et al.
Pubblicazione: (2025)
di: Romero, David, et al.
Pubblicazione: (2025)
InterRVOS: Interaction-aware Referring Video Object Segmentation
di: Jin, Woojeong, et al.
Pubblicazione: (2025)
di: Jin, Woojeong, et al.
Pubblicazione: (2025)
HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation
di: Liu, Kun, et al.
Pubblicazione: (2025)
di: Liu, Kun, et al.
Pubblicazione: (2025)
MotionFix: Text-Driven 3D Human Motion Editing
di: Athanasiou, Nikos, et al.
Pubblicazione: (2024)
di: Athanasiou, Nikos, et al.
Pubblicazione: (2024)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
di: Chen, Zerui, et al.
Pubblicazione: (2024)
di: Chen, Zerui, et al.
Pubblicazione: (2024)
Mitigating Object Hallucination via Concentric Causal Attention
di: Xing, Yun, et al.
Pubblicazione: (2024)
di: Xing, Yun, et al.
Pubblicazione: (2024)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
di: Han, Mingfei, et al.
Pubblicazione: (2026)
di: Han, Mingfei, et al.
Pubblicazione: (2026)
Learning Human-Object Interaction for 3D Human Pose Estimation from LiDAR Point Clouds
di: Jung, Daniel Sungho, et al.
Pubblicazione: (2026)
di: Jung, Daniel Sungho, et al.
Pubblicazione: (2026)
HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
di: Huang, Ziyao, et al.
Pubblicazione: (2025)
di: Huang, Ziyao, et al.
Pubblicazione: (2025)
Learning to Generate Human-Human-Object Interactions from Textual Descriptions
di: Na, Jeonghyeon, et al.
Pubblicazione: (2025)
di: Na, Jeonghyeon, et al.
Pubblicazione: (2025)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
di: Zhou, Donghao, et al.
Pubblicazione: (2026)
di: Zhou, Donghao, et al.
Pubblicazione: (2026)
PhysiInter: Integrating Physical Mapping for High-Fidelity Human Interaction Generation
di: Yao, Wei, et al.
Pubblicazione: (2025)
di: Yao, Wei, et al.
Pubblicazione: (2025)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
di: Jang, Youngjoon, et al.
Pubblicazione: (2025)
di: Jang, Youngjoon, et al.
Pubblicazione: (2025)
UnrealPose: Leveraging Game Engine Kinematics for Large-Scale Synthetic Human Pose Data
di: Kawaguchi, Joshua, et al.
Pubblicazione: (2026)
di: Kawaguchi, Joshua, et al.
Pubblicazione: (2026)
PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
di: He, Jingxuan, et al.
Pubblicazione: (2025)
di: He, Jingxuan, et al.
Pubblicazione: (2025)
InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
di: Wu, Zizhao, et al.
Pubblicazione: (2025)
di: Wu, Zizhao, et al.
Pubblicazione: (2025)
VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
di: Zhang, Wanyue, et al.
Pubblicazione: (2025)
di: Zhang, Wanyue, et al.
Pubblicazione: (2025)
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
di: Gan, Qijun, et al.
Pubblicazione: (2025)
di: Gan, Qijun, et al.
Pubblicazione: (2025)
Learning Human-Object Interaction as Groups
di: Hong, Jiajun, et al.
Pubblicazione: (2025)
di: Hong, Jiajun, et al.
Pubblicazione: (2025)
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
di: Luo, Xiangyang, et al.
Pubblicazione: (2026)
di: Luo, Xiangyang, et al.
Pubblicazione: (2026)
InTraGen: Trajectory-controlled Video Generation for Object Interactions
di: Liu, Zuhao, et al.
Pubblicazione: (2024)
di: Liu, Zuhao, et al.
Pubblicazione: (2024)
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
di: Fan, Sicheng, et al.
Pubblicazione: (2026)
di: Fan, Sicheng, et al.
Pubblicazione: (2026)
AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos
di: Lu, Feichi, et al.
Pubblicazione: (2024)
di: Lu, Feichi, et al.
Pubblicazione: (2024)
Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
di: Fan, Zicong, et al.
Pubblicazione: (2024)
di: Fan, Zicong, et al.
Pubblicazione: (2024)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
di: Ma, Yue, et al.
Pubblicazione: (2023)
di: Ma, Yue, et al.
Pubblicazione: (2023)
Text-Driven 3D Hand Motion Generation from Sign Language Data
di: Bensabath, Léore, et al.
Pubblicazione: (2025)
di: Bensabath, Léore, et al.
Pubblicazione: (2025)
InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
di: Wang, Zhenzhi, et al.
Pubblicazione: (2023)
di: Wang, Zhenzhi, et al.
Pubblicazione: (2023)
Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation
di: Petrovich, Mathis, et al.
Pubblicazione: (2024)
di: Petrovich, Mathis, et al.
Pubblicazione: (2024)
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
di: Jiao, Yingying, et al.
Pubblicazione: (2025)
di: Jiao, Yingying, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
di: Zhang, Yangsong, et al.
Pubblicazione: (2026) -
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
di: Xu, Sirui, et al.
Pubblicazione: (2025) -
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
di: Bensabath, Léore, et al.
Pubblicazione: (2024) -
SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation
di: Athanasiou, Nikos, et al.
Pubblicazione: (2023) -
Learning text-to-video retrieval from image captioning
di: Ventura, Lucas, et al.
Pubblicazione: (2024)