Multi-speaker Attention Alignment for Multimodal Social Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Ouyang, Liangyang, Huang, Yifei, Zhang, Mingfang, Kang, Caixin, Furuta, Ryosuke, Sato, Yoichi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025)
by: Kang, Caixin, et al.
Published: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025)
by: Kang, Caixin, et al.
Published: (2025)
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
by: Ouyang, Liangyang, et al.
Published: (2026)
by: Ouyang, Liangyang, et al.
Published: (2026)
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
by: Liu, Ruicong, et al.
Published: (2025)
by: Liu, Ruicong, et al.
Published: (2025)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
by: Zhang, Mingfang, et al.
Published: (2025)
by: Zhang, Mingfang, et al.
Published: (2025)
Leadership Assessment in Pediatric Intensive Care Unit Team Training
by: Ouyang, Liangyang, et al.
Published: (2025)
by: Ouyang, Liangyang, et al.
Published: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
by: Lin, Nie, et al.
Published: (2024)
by: Lin, Nie, et al.
Published: (2024)
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
by: Furuta, Ryosuke, et al.
Published: (2023)
by: Furuta, Ryosuke, et al.
Published: (2023)
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
by: Kang, Caixin, et al.
Published: (2026)
by: Kang, Caixin, et al.
Published: (2026)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
by: Sakai, Yuki, et al.
Published: (2025)
by: Sakai, Yuki, et al.
Published: (2025)
SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training
by: Lin, Nie, et al.
Published: (2025)
by: Lin, Nie, et al.
Published: (2025)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
Learning Multiple Object States from Actions via Large Language Models
by: Tateno, Masatoshi, et al.
Published: (2024)
by: Tateno, Masatoshi, et al.
Published: (2024)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
by: Zhang, Mingfang, et al.
Published: (2026)
by: Zhang, Mingfang, et al.
Published: (2026)
FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
by: Yagi, Takuma, et al.
Published: (2024)
by: Yagi, Takuma, et al.
Published: (2024)
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
by: Banno, Tatsuro, et al.
Published: (2025)
by: Banno, Tatsuro, et al.
Published: (2025)
LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
by: Ouyang, Liangyang, et al.
Published: (2025)
by: Ouyang, Liangyang, et al.
Published: (2025)
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
by: Liu, Ruicong, et al.
Published: (2024)
by: Liu, Ruicong, et al.
Published: (2024)
Inference-time Trajectory Optimization for Manga Image Editing
by: Furuta, Ryosuke
Published: (2026)
by: Furuta, Ryosuke
Published: (2026)
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
by: Suzuki, Naru, et al.
Published: (2025)
by: Suzuki, Naru, et al.
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
by: Taniguchi, Takara, et al.
Published: (2024)
by: Taniguchi, Takara, et al.
Published: (2024)
Egocentric Gaze Estimation via Neck-Mounted Camera
by: Huang, Haoyu, et al.
Published: (2026)
by: Huang, Haoyu, et al.
Published: (2026)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
by: Liu, Ruicong, et al.
Published: (2025)
by: Liu, Ruicong, et al.
Published: (2025)
Generative Modeling of Shape-Dependent Self-Contact Human Poses
by: Ohkawa, Takehiko, et al.
Published: (2025)
by: Ohkawa, Takehiko, et al.
Published: (2025)
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
by: Lin, Nie, et al.
Published: (2025)
by: Lin, Nie, et al.
Published: (2025)
Interactive Multi-Head Self-Attention with Linear Complexity
by: Kang, Hankyul, et al.
Published: (2024)
by: Kang, Hankyul, et al.
Published: (2024)
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
by: Tateno, Masatoshi, et al.
Published: (2025)
by: Tateno, Masatoshi, et al.
Published: (2025)
Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
by: Kang, Donghwa, et al.
Published: (2025)
by: Kang, Donghwa, et al.
Published: (2025)
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
by: Ruan, Shouwei, et al.
Published: (2025)
by: Ruan, Shouwei, et al.
Published: (2025)
Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
by: Park, Jonghyun, et al.
Published: (2025)
by: Park, Jonghyun, et al.
Published: (2025)
PSDiffusion: Harmonized Multi-Layer Image Generation via Layout and Appearance Alignment
by: Huang, Dingbang, et al.
Published: (2025)
by: Huang, Dingbang, et al.
Published: (2025)
AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?
by: Ruan, Shouwei, et al.
Published: (2024)
by: Ruan, Shouwei, et al.
Published: (2024)
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
by: Kwon, Minkyung, et al.
Published: (2025)
by: Kwon, Minkyung, et al.
Published: (2025)
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
Similar Items
-
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025) -
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025) -
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
by: Ouyang, Liangyang, et al.
Published: (2026) -
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
by: Liu, Ruicong, et al.
Published: (2025) -
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)