Populate-A-Scene: Affordance-Aware Human Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shan, Mengyi, He, Zecheng, Ma, Haoyu, Juefei-Xu, Felix, Zhang, Peizhao, Hou, Tingbo, Chuang, Ching-Yao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
von: Song, Kunpeng, et al.
Veröffentlicht: (2024)
von: Song, Kunpeng, et al.
Veröffentlicht: (2024)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
von: Liang, Feng, et al.
Veröffentlicht: (2025)
von: Liang, Feng, et al.
Veröffentlicht: (2025)
MoCha: Towards Movie-Grade Talking Character Synthesis
von: Wei, Cong, et al.
Veröffentlicht: (2025)
von: Wei, Cong, et al.
Veröffentlicht: (2025)
StreamDiT: Real-Time Streaming Text-to-Video Generation
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
von: Ma, Xu, et al.
Veröffentlicht: (2025)
von: Ma, Xu, et al.
Veröffentlicht: (2025)
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
von: Chen, Chao, et al.
Veröffentlicht: (2023)
von: Chen, Chao, et al.
Veröffentlicht: (2023)
Imagine yourself: Tuning-Free Personalized Image Generation
von: He, Zecheng, et al.
Veröffentlicht: (2024)
von: He, Zecheng, et al.
Veröffentlicht: (2024)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
von: Wang, Qirui, et al.
Veröffentlicht: (2026)
von: Wang, Qirui, et al.
Veröffentlicht: (2026)
MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
von: Zhao, Yang, et al.
Veröffentlicht: (2023)
von: Zhao, Yang, et al.
Veröffentlicht: (2023)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
von: Mao, Aihua, et al.
Veröffentlicht: (2026)
von: Mao, Aihua, et al.
Veröffentlicht: (2026)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
von: Wang, Zan, et al.
Veröffentlicht: (2024)
von: Wang, Zan, et al.
Veröffentlicht: (2024)
Instance Tracking in 3D Scenes from Egocentric Videos
von: Zhao, Yunhan, et al.
Veröffentlicht: (2023)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2023)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
Hazard-Aware Traffic Scene Graph Generation
von: Huang, Yaoqi, et al.
Veröffentlicht: (2026)
von: Huang, Yaoqi, et al.
Veröffentlicht: (2026)
TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances
von: Xu, Wenting, et al.
Veröffentlicht: (2024)
von: Xu, Wenting, et al.
Veröffentlicht: (2024)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion
von: He, Jixuan, et al.
Veröffentlicht: (2024)
von: He, Jixuan, et al.
Veröffentlicht: (2024)
$α$-OCC: Uncertainty-Aware Camera-based 3D Semantic Occupancy Prediction
von: Su, Sanbao, et al.
Veröffentlicht: (2024)
von: Su, Sanbao, et al.
Veröffentlicht: (2024)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
von: Wu, Xiaofei, et al.
Veröffentlicht: (2026)
von: Wu, Xiaofei, et al.
Veröffentlicht: (2026)
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
von: Maillard, Léopold, et al.
Veröffentlicht: (2026)
von: Maillard, Léopold, et al.
Veröffentlicht: (2026)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
von: Zhao, Zecheng, et al.
Veröffentlicht: (2025)
von: Zhao, Zecheng, et al.
Veröffentlicht: (2025)
Human-Aware 3D Scene Generation with Spatially-constrained Diffusion Models
von: Hong, Xiaolin, et al.
Veröffentlicht: (2024)
von: Hong, Xiaolin, et al.
Veröffentlicht: (2024)
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
von: Wang, Hanqing, et al.
Veröffentlicht: (2026)
von: Wang, Hanqing, et al.
Veröffentlicht: (2026)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
von: Mei, Jianbiao, et al.
Veröffentlicht: (2024)
von: Mei, Jianbiao, et al.
Veröffentlicht: (2024)
GenHSI: Controllable Generation of Human-Scene Interaction Videos
von: Li, Zekun, et al.
Veröffentlicht: (2025)
von: Li, Zekun, et al.
Veröffentlicht: (2025)
Affordance-First Decomposition for Continual Learning in Video-Language Understanding
von: Xu, Mengzhu, et al.
Veröffentlicht: (2025)
von: Xu, Mengzhu, et al.
Veröffentlicht: (2025)
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
von: Fang, Zhixue, et al.
Veröffentlicht: (2026)
von: Fang, Zhixue, et al.
Veröffentlicht: (2026)
INGeo: Accelerating Instant Neural Scene Reconstruction with Noisy Geometry Priors
von: Li, Chaojian, et al.
Veröffentlicht: (2022)
von: Li, Chaojian, et al.
Veröffentlicht: (2022)
Text-Pass Filter: An Efficient Scene Text Detector
von: Yang, Chuang, et al.
Veröffentlicht: (2026)
von: Yang, Chuang, et al.
Veröffentlicht: (2026)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
von: Yao, Wei, et al.
Veröffentlicht: (2025)
von: Yao, Wei, et al.
Veröffentlicht: (2025)
Physics-Aware 3D Gaussian Editing for Driving Scene Generation
von: Zhou, Feng, et al.
Veröffentlicht: (2026)
von: Zhou, Feng, et al.
Veröffentlicht: (2026)
AMG: Avatar Motion Guided Video Generation
von: Yang, Zhangsihao, et al.
Veröffentlicht: (2024)
von: Yang, Zhangsihao, et al.
Veröffentlicht: (2024)
3D Congealing: 3D-Aware Image Alignment in the Wild
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
von: Zhou, Dingyi, et al.
Veröffentlicht: (2026)
von: Zhou, Dingyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
von: Song, Kunpeng, et al.
Veröffentlicht: (2024) -
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
von: Zhang, Haochen, et al.
Veröffentlicht: (2026) -
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
von: Liang, Feng, et al.
Veröffentlicht: (2025) -
MoCha: Towards Movie-Grade Talking Character Synthesis
von: Wei, Cong, et al.
Veröffentlicht: (2025) -
StreamDiT: Real-Time Streaming Text-to-Video Generation
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)