Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Sangmin, Lai, Bolin, Ryan, Fiona, Boote, Bikram, Rehg, James M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
von: Kim, Junho, et al.
Veröffentlicht: (2026)
von: Kim, Junho, et al.
Veröffentlicht: (2026)
SocialGesture: Delving into Multi-person Gesture Understanding
von: Cao, Xu, et al.
Veröffentlicht: (2025)
von: Cao, Xu, et al.
Veröffentlicht: (2025)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
von: Lai, Bolin, et al.
Veröffentlicht: (2022)
von: Lai, Bolin, et al.
Veröffentlicht: (2022)
MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models
von: Thai, Anh, et al.
Veröffentlicht: (2025)
von: Thai, Anh, et al.
Veröffentlicht: (2025)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
Leveraging Object Priors for Point Tracking
von: Boote, Bikram, et al.
Veröffentlicht: (2024)
von: Boote, Bikram, et al.
Veröffentlicht: (2024)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
Towards Online Multi-Modal Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders
von: Ryan, Fiona, et al.
Veröffentlicht: (2024)
von: Ryan, Fiona, et al.
Veröffentlicht: (2024)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
Towards Social AI: A Survey on Understanding Social Interactions
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
Learning Predictive Visuomotor Coordination
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Bilingual Text-to-Motion Generation: A New Benchmark and Baselines
von: Weng, Wanjiang, et al.
Veröffentlicht: (2026)
von: Weng, Wanjiang, et al.
Veröffentlicht: (2026)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
von: Yuxuan, Cao, et al.
Veröffentlicht: (2025)
von: Yuxuan, Cao, et al.
Veröffentlicht: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
GLOS: Sign Language Generation with Temporally Aligned Gloss-Level Conditioning
von: Lee, Taeryung, et al.
Veröffentlicht: (2025)
von: Lee, Taeryung, et al.
Veröffentlicht: (2025)
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
von: Wang, Chuhan, et al.
Veröffentlicht: (2026)
von: Wang, Chuhan, et al.
Veröffentlicht: (2026)
Improving Personalized Search with Regularized Low-Rank Parameter Updates
von: Ryan, Fiona, et al.
Veröffentlicht: (2025)
von: Ryan, Fiona, et al.
Veröffentlicht: (2025)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
Toward Interactive Regional Understanding in Vision-Large Language Models
von: Lee, Jungbeom, et al.
Veröffentlicht: (2024)
von: Lee, Jungbeom, et al.
Veröffentlicht: (2024)
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
von: Huang, Runsheng, et al.
Veröffentlicht: (2024)
von: Huang, Runsheng, et al.
Veröffentlicht: (2024)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
Do Multimodal Large Language Models Understand Welding?
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
von: Toschi, Federico, et al.
Veröffentlicht: (2026)
von: Toschi, Federico, et al.
Veröffentlicht: (2026)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
von: Huang, Kui, et al.
Veröffentlicht: (2025)
von: Huang, Kui, et al.
Veröffentlicht: (2025)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
von: Kim, Junho, et al.
Veröffentlicht: (2026) -
SocialGesture: Delving into Multi-person Gesture Understanding
von: Cao, Xu, et al.
Veröffentlicht: (2025) -
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
von: Lai, Bolin, et al.
Veröffentlicht: (2022) -
MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models
von: Thai, Anh, et al.
Veröffentlicht: (2025) -
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
von: Lai, Bolin, et al.
Veröffentlicht: (2023)