GenHSI: Controllable Generation of Human-Scene Interaction Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zekun, Zhou, Rui, Sajnani, Rahul, Cong, Xiaoyan, Ritchie, Daniel, Sridhar, Srinath |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoDiffuser: Geometry-Based Image Editing with Diffusion Models
by: Sajnani, Rahul, et al.
Published: (2024)
by: Sajnani, Rahul, et al.
Published: (2024)
GenHeld: Generating and Editing Handheld Objects
by: Min, Chaerin, et al.
Published: (2024)
by: Min, Chaerin, et al.
Published: (2024)
Art3D: Training-Free 3D Generation from Flat-Colored Illustration
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation
by: Li, Hongjie, et al.
Published: (2024)
by: Li, Hongjie, et al.
Published: (2024)
CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes
by: Xu, Xianghao, et al.
Published: (2024)
by: Xu, Xianghao, et al.
Published: (2024)
PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
by: Rai, Aashish, et al.
Published: (2026)
by: Rai, Aashish, et al.
Published: (2026)
DyTact: Capturing Dynamic Contacts in Hand-Object Manipulation
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
by: Pan, Liang, et al.
Published: (2025)
by: Pan, Liang, et al.
Published: (2025)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
by: Rai, Aashish, et al.
Published: (2024)
by: Rai, Aashish, et al.
Published: (2024)
FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework
by: Mu, Lingzhou, et al.
Published: (2025)
by: Mu, Lingzhou, et al.
Published: (2025)
GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities
by: Fu, Rao, et al.
Published: (2024)
by: Fu, Rao, et al.
Published: (2024)
Dropout Concrete Autoencoder for Band Selection on HSI Scenes
by: Xu, Lei, et al.
Published: (2024)
by: Xu, Lei, et al.
Published: (2024)
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
by: Cong, Xiaoyan, et al.
Published: (2026)
by: Cong, Xiaoyan, et al.
Published: (2026)
GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?
by: Zou, Yueying, et al.
Published: (2026)
by: Zou, Yueying, et al.
Published: (2026)
InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable Gaussians
by: Chen, Kefan, et al.
Published: (2025)
by: Chen, Kefan, et al.
Published: (2025)
FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
by: Chen, Kefan, et al.
Published: (2024)
by: Chen, Kefan, et al.
Published: (2024)
MANUS: Markerless Grasp Capture using Articulated 3D Gaussians
by: Pokhariya, Chandradeep, et al.
Published: (2023)
by: Pokhariya, Chandradeep, et al.
Published: (2023)
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
by: Li, Zekun, et al.
Published: (2026)
by: Li, Zekun, et al.
Published: (2026)
Generating Human Interaction Motions in Scenes with Text Control
by: Yi, Hongwei, et al.
Published: (2024)
by: Yi, Hongwei, et al.
Published: (2024)
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
by: Zhai, Shangjin, et al.
Published: (2025)
by: Zhai, Shangjin, et al.
Published: (2025)
SimGen: Simulator-conditioned Driving Scene Generation
by: Zhou, Yunsong, et al.
Published: (2024)
by: Zhou, Yunsong, et al.
Published: (2024)
Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes
by: Wang, Yin, et al.
Published: (2026)
by: Wang, Yin, et al.
Published: (2026)
PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
by: He, Jingxuan, et al.
Published: (2025)
by: He, Jingxuan, et al.
Published: (2025)
DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
LaGen: Towards Autoregressive LiDAR Scene Generation
by: Zhou, Sizhuo, et al.
Published: (2025)
by: Zhou, Sizhuo, et al.
Published: (2025)
SceneMI: Motion In-betweening for Modeling Human-Scene Interactions
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
GenSelfDiff-HIS: Generative Self-Supervision Using Diffusion for Histopathological Image Segmentation
by: Purma, Vishnuvardhan, et al.
Published: (2023)
by: Purma, Vishnuvardhan, et al.
Published: (2023)
InTraGen: Trajectory-controlled Video Generation for Object Interactions
by: Liu, Zuhao, et al.
Published: (2024)
by: Liu, Zuhao, et al.
Published: (2024)
AnyHome: Open-Vocabulary Generation of Structured and Textured 3D Homes
by: Fu, Rao, et al.
Published: (2023)
by: Fu, Rao, et al.
Published: (2023)
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
by: Du, Hongyang, et al.
Published: (2026)
by: Du, Hongyang, et al.
Published: (2026)
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023)
by: Chen, Shoufa, et al.
Published: (2023)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Electrolyzers-HSI: Close-Range Multi-Scene Hyperspectral Imaging Benchmark Dataset
by: Arbash, Elias, et al.
Published: (2025)
by: Arbash, Elias, et al.
Published: (2025)
WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation
by: Lu, Jiachen, et al.
Published: (2023)
by: Lu, Jiachen, et al.
Published: (2023)
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
by: Li, Runjia, et al.
Published: (2025)
by: Li, Runjia, et al.
Published: (2025)
PhysiGen: Integrating Collision-Aware Physical Constraints for High-Fidelity Human-Human Interaction Generation
by: Lei, Nan, et al.
Published: (2026)
by: Lei, Nan, et al.
Published: (2026)
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
by: Lee, JoungBin, et al.
Published: (2025)
by: Lee, JoungBin, et al.
Published: (2025)
R3DS: Reality-linked 3D Scenes for Panoramic Scene Understanding
by: Wu, Qirui, et al.
Published: (2024)
by: Wu, Qirui, et al.
Published: (2024)
Populate-A-Scene: Affordance-Aware Human Video Generation
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
Similar Items
-
GeoDiffuser: Geometry-Based Image Editing with Diffusion Models
by: Sajnani, Rahul, et al.
Published: (2024) -
GenHeld: Generating and Editing Handheld Objects
by: Min, Chaerin, et al.
Published: (2024) -
Art3D: Training-Free 3D Generation from Flat-Colored Illustration
by: Cong, Xiaoyan, et al.
Published: (2025) -
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation
by: Li, Hongjie, et al.
Published: (2024) -
CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes
by: Xu, Xianghao, et al.
Published: (2024)