Towards Online Multi-Modal Social Interaction Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinpeng, Deng, Shijian, Lai, Bolin, Pian, Weiguo, Rehg, James M., Tian, Yapeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2026)
by: Li, Xinpeng, et al.
Published: (2026)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
by: Pian, Weiguo, et al.
Published: (2026)
by: Pian, Weiguo, et al.
Published: (2026)
SocialGesture: Delving into Multi-person Gesture Understanding
by: Cao, Xu, et al.
Published: (2025)
by: Cao, Xu, et al.
Published: (2025)
Continual Audio-Visual Sound Separation
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
by: Lee, Sangmin, et al.
Published: (2024)
by: Lee, Sangmin, et al.
Published: (2024)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
by: Lai, Bolin, et al.
Published: (2022)
by: Lai, Bolin, et al.
Published: (2022)
Do Joint Audio-Video Generation Models Understand Physics?
by: Cui, Zijun, et al.
Published: (2026)
by: Cui, Zijun, et al.
Published: (2026)
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Learning Predictive Visuomotor Coordination
by: Jia, Wenqi, et al.
Published: (2025)
by: Jia, Wenqi, et al.
Published: (2025)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
by: Tian, Huiyuan, et al.
Published: (2025)
by: Tian, Huiyuan, et al.
Published: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
by: Ding, Xinpeng, et al.
Published: (2024)
by: Ding, Xinpeng, et al.
Published: (2024)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
Explainable AI-Generated Image Detection RewardBench
by: Yang, Michael, et al.
Published: (2025)
by: Yang, Michael, et al.
Published: (2025)
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models
by: Thai, Anh, et al.
Published: (2025)
by: Thai, Anh, et al.
Published: (2025)
Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Cross-Modal Learning for Anomaly Detection in Complex Industrial Process: Methodology and Benchmark
by: Wu, Gaochang, et al.
Published: (2024)
by: Wu, Gaochang, et al.
Published: (2024)
DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Vinedresser3D: Agentic Text-guided 3D Editing
by: Chi, Yankuan, et al.
Published: (2026)
by: Chi, Yankuan, et al.
Published: (2026)
Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
DeepInteraction++: Multi-Modality Interaction for Autonomous Driving
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
Toward Gaze Target Detection of Young Autistic Children
by: Deng, Shijian, et al.
Published: (2025)
by: Deng, Shijian, et al.
Published: (2025)
From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
by: Tian, Huiyuan, et al.
Published: (2025)
by: Tian, Huiyuan, et al.
Published: (2025)
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
by: Deng, Shijian, et al.
Published: (2024)
by: Deng, Shijian, et al.
Published: (2024)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
MMGait: Towards Multi-Modal Gait Recognition
by: Wang, Chenye, et al.
Published: (2026)
by: Wang, Chenye, et al.
Published: (2026)
Human Action Anticipation: A Survey
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
Segment Anything with Multiple Modalities
by: Xiao, Aoran, et al.
Published: (2024)
by: Xiao, Aoran, et al.
Published: (2024)
Similar Items
-
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2026) -
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
by: Pian, Weiguo, et al.
Published: (2024) -
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
by: Pian, Weiguo, et al.
Published: (2026) -
SocialGesture: Delving into Multi-person Gesture Understanding
by: Cao, Xu, et al.
Published: (2025) -
Continual Audio-Visual Sound Separation
by: Pian, Weiguo, et al.
Published: (2024)