SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Ouyang, Liangyang, Liu, Ruicong, Kang, Caixin, Huang, Yifei, Sato, Yoichi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
por: Kang, Caixin, et al.
Publicado: (2025)
por: Kang, Caixin, et al.
Publicado: (2025)
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
por: Liu, Ruicong, et al.
Publicado: (2025)
por: Liu, Ruicong, et al.
Publicado: (2025)
Multi-speaker Attention Alignment for Multimodal Social Interaction
por: Ouyang, Liangyang, et al.
Publicado: (2025)
por: Ouyang, Liangyang, et al.
Publicado: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
por: Kang, Caixin, et al.
Publicado: (2025)
por: Kang, Caixin, et al.
Publicado: (2025)
ActionVOS: Actions as Prompts for Video Object Segmentation
por: Ouyang, Liangyang, et al.
Publicado: (2024)
por: Ouyang, Liangyang, et al.
Publicado: (2024)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
por: Zhang, Mingfang, et al.
Publicado: (2025)
por: Zhang, Mingfang, et al.
Publicado: (2025)
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
por: Kang, Caixin, et al.
Publicado: (2026)
por: Kang, Caixin, et al.
Publicado: (2026)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
por: Zhang, Mingfang, et al.
Publicado: (2024)
por: Zhang, Mingfang, et al.
Publicado: (2024)
Leadership Assessment in Pediatric Intensive Care Unit Team Training
por: Ouyang, Liangyang, et al.
Publicado: (2025)
por: Ouyang, Liangyang, et al.
Publicado: (2025)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
por: Zhu, Zhifan, et al.
Publicado: (2025)
por: Zhu, Zhifan, et al.
Publicado: (2025)
LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
por: Ouyang, Liangyang, et al.
Publicado: (2025)
por: Ouyang, Liangyang, et al.
Publicado: (2025)
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
por: Liu, Ruicong, et al.
Publicado: (2024)
por: Liu, Ruicong, et al.
Publicado: (2024)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
por: Zhang, Mingfang, et al.
Publicado: (2026)
por: Zhang, Mingfang, et al.
Publicado: (2026)
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
por: Banno, Tatsuro, et al.
Publicado: (2025)
por: Banno, Tatsuro, et al.
Publicado: (2025)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
por: Liu, Ruicong, et al.
Publicado: (2025)
por: Liu, Ruicong, et al.
Publicado: (2025)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
por: Sakai, Yuki, et al.
Publicado: (2025)
por: Sakai, Yuki, et al.
Publicado: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
por: Lin, Nie, et al.
Publicado: (2024)
por: Lin, Nie, et al.
Publicado: (2024)
DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
por: Peng, Yichen, et al.
Publicado: (2026)
por: Peng, Yichen, et al.
Publicado: (2026)
FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
por: Yagi, Takuma, et al.
Publicado: (2024)
por: Yagi, Takuma, et al.
Publicado: (2024)
SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
por: Xu, Qianxun, et al.
Publicado: (2026)
por: Xu, Qianxun, et al.
Publicado: (2026)
Egocentric Gaze Estimation via Neck-Mounted Camera
por: Huang, Haoyu, et al.
Publicado: (2026)
por: Huang, Haoyu, et al.
Publicado: (2026)
ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
por: Wu, Xiaoxue, et al.
Publicado: (2025)
por: Wu, Xiaoxue, et al.
Publicado: (2025)
FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
por: Li, Guangzhao, et al.
Publicado: (2025)
por: Li, Guangzhao, et al.
Publicado: (2025)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
por: Chen, Junhao, et al.
Publicado: (2025)
por: Chen, Junhao, et al.
Publicado: (2025)
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
por: Furuta, Ryosuke, et al.
Publicado: (2023)
por: Furuta, Ryosuke, et al.
Publicado: (2023)
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
por: Tateno, Masatoshi, et al.
Publicado: (2025)
por: Tateno, Masatoshi, et al.
Publicado: (2025)
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
por: Zhong, Zhizhou, et al.
Publicado: (2025)
por: Zhong, Zhizhou, et al.
Publicado: (2025)
Training-Free Robust Interactive Video Object Segmentation
por: Wei, Xiaoli, et al.
Publicado: (2024)
por: Wei, Xiaoli, et al.
Publicado: (2024)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation
por: Ling, Pengyang, et al.
Publicado: (2024)
por: Ling, Pengyang, et al.
Publicado: (2024)
UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
por: Chu, Xuangeng, et al.
Publicado: (2025)
por: Chu, Xuangeng, et al.
Publicado: (2025)
Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image Generation
por: Huang, Qihan, et al.
Publicado: (2024)
por: Huang, Qihan, et al.
Publicado: (2024)
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
por: Lu, Yu, et al.
Publicado: (2025)
por: Lu, Yu, et al.
Publicado: (2025)
Towards Interactive Intelligence for Digital Humans
por: Cai, Yiyi, et al.
Publicado: (2025)
por: Cai, Yiyi, et al.
Publicado: (2025)
VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree
por: Li, Wenlong, et al.
Publicado: (2025)
por: Li, Wenlong, et al.
Publicado: (2025)
LongDiff: Training-Free Long Video Generation in One Go
por: Li, Zhuoling, et al.
Publicado: (2025)
por: Li, Zhuoling, et al.
Publicado: (2025)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
por: Mazzamuto, Michele, et al.
Publicado: (2024)
por: Mazzamuto, Michele, et al.
Publicado: (2024)
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
por: Kong, Zhe, et al.
Publicado: (2025)
por: Kong, Zhe, et al.
Publicado: (2025)
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
por: Feng, X., et al.
Publicado: (2026)
por: Feng, X., et al.
Publicado: (2026)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
por: Peng, Bo, et al.
Publicado: (2023)
por: Peng, Bo, et al.
Publicado: (2023)
SocialGen: Modeling Multi-Human Social Interaction with Language Models
por: Yu, Heng, et al.
Publicado: (2025)
por: Yu, Heng, et al.
Publicado: (2025)
Ejemplares similares
-
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
por: Kang, Caixin, et al.
Publicado: (2025) -
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
por: Liu, Ruicong, et al.
Publicado: (2025) -
Multi-speaker Attention Alignment for Multimodal Social Interaction
por: Ouyang, Liangyang, et al.
Publicado: (2025) -
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
por: Kang, Caixin, et al.
Publicado: (2025) -
ActionVOS: Actions as Prompts for Video Object Segmentation
por: Ouyang, Liangyang, et al.
Publicado: (2024)