GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Junho, Cao, Xu, Yang, Houze, Boote, Bikram, Jojic, Ana, Ryan, Fiona, Lai, Bolin, Lee, Sangmin, Rehg, James M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
Towards Social AI: A Survey on Understanding Social Interactions
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
SocialGesture: Delving into Multi-person Gesture Understanding
von: Cao, Xu, et al.
Veröffentlicht: (2025)
von: Cao, Xu, et al.
Veröffentlicht: (2025)
Leveraging Object Priors for Point Tracking
von: Boote, Bikram, et al.
Veröffentlicht: (2024)
von: Boote, Bikram, et al.
Veröffentlicht: (2024)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
von: Lai, Bolin, et al.
Veröffentlicht: (2022)
von: Lai, Bolin, et al.
Veröffentlicht: (2022)
MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models
von: Thai, Anh, et al.
Veröffentlicht: (2025)
von: Thai, Anh, et al.
Veröffentlicht: (2025)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
Towards Online Multi-Modal Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders
von: Ryan, Fiona, et al.
Veröffentlicht: (2024)
von: Ryan, Fiona, et al.
Veröffentlicht: (2024)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
von: Kim, Junho, et al.
Veröffentlicht: (2026)
von: Kim, Junho, et al.
Veröffentlicht: (2026)
Improving Personalized Search with Regularized Low-Rank Parameter Updates
von: Ryan, Fiona, et al.
Veröffentlicht: (2025)
von: Ryan, Fiona, et al.
Veröffentlicht: (2025)
VG3T: Visual Geometry Grounded Gaussian Transformer
von: Kim, Junho, et al.
Veröffentlicht: (2025)
von: Kim, Junho, et al.
Veröffentlicht: (2025)
GenZ: Foundational models as latent variable generators within traditional statistical models
von: Jojic, Marko, et al.
Veröffentlicht: (2025)
von: Jojic, Marko, et al.
Veröffentlicht: (2025)
GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification
von: Wan, Faxian, et al.
Veröffentlicht: (2026)
von: Wan, Faxian, et al.
Veröffentlicht: (2026)
C-GRASP: Clinically-Grounded Reasoning for Affective Signal Processing
von: Cheng, Cheng Lin, et al.
Veröffentlicht: (2026)
von: Cheng, Cheng Lin, et al.
Veröffentlicht: (2026)
Learning Predictive Visuomotor Coordination
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
von: Lee, Hojae, et al.
Veröffentlicht: (2024)
von: Lee, Hojae, et al.
Veröffentlicht: (2024)
Learning Informative Attention Weights for Person Re-Identification
von: Wang, Yancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yancheng, et al.
Veröffentlicht: (2025)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
SoftEDA: Rethinking Rule-Based Data Augmentation with Soft Labels
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
AutoAugment Is What You Need: Enhancing Rule-based Augmentation Methods in Low-resource Regimes
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
von: Jin, Kyohoon, et al.
Veröffentlicht: (2024)
von: Jin, Kyohoon, et al.
Veröffentlicht: (2024)
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
von: Kim, Hyung Kyu, et al.
Veröffentlicht: (2025)
von: Kim, Hyung Kyu, et al.
Veröffentlicht: (2025)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
Narrative-Driven Paper-to-Slide Generation via ArcDeck
von: Ozden, Tarik Can, et al.
Veröffentlicht: (2026)
von: Ozden, Tarik Can, et al.
Veröffentlicht: (2026)
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
von: Kim, Jihyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jihyeon, et al.
Veröffentlicht: (2026)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
On the integrability of root-Kerr probe dynamics
von: Kim, Sungsoo, et al.
Veröffentlicht: (2026)
von: Kim, Sungsoo, et al.
Veröffentlicht: (2026)
Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
von: Kim, Sangmin, et al.
Veröffentlicht: (2026)
von: Kim, Sangmin, et al.
Veröffentlicht: (2026)
GRASP: group-Shapley feature selection for patients
von: Luo, Yuheng, et al.
Veröffentlicht: (2026)
von: Luo, Yuheng, et al.
Veröffentlicht: (2026)
Sperner's colorings of hypergraphs arising from edgewise triangulations
von: Jojić, Duško, et al.
Veröffentlicht: (2025)
von: Jojić, Duško, et al.
Veröffentlicht: (2025)
A labeling of the Simplex-Lattice Hypergraph with at most 2 colors on each hyperedge
von: Papaz, Ognjen, et al.
Veröffentlicht: (2025)
von: Papaz, Ognjen, et al.
Veröffentlicht: (2025)
Shelling of links and star clusters in edgewise subdivision of a simplex
von: Jojić, Duško, et al.
Veröffentlicht: (2024)
von: Jojić, Duško, et al.
Veröffentlicht: (2024)
Stability of political structures modeled by simplicial complexes under mediation, splitting, and shellability
von: Jojić, Duško, et al.
Veröffentlicht: (2025)
von: Jojić, Duško, et al.
Veröffentlicht: (2025)
Non-freeness of parabolic two-generator groups
von: Choi, Philip, et al.
Veröffentlicht: (2024)
von: Choi, Philip, et al.
Veröffentlicht: (2024)
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
von: Jang, Chaeyun, et al.
Veröffentlicht: (2025)
von: Jang, Chaeyun, et al.
Veröffentlicht: (2025)
OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning
von: Bi, Zhenyu, et al.
Veröffentlicht: (2025)
von: Bi, Zhenyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
von: Lee, Sangmin, et al.
Veröffentlicht: (2024) -
Towards Social AI: A Survey on Understanding Social Interactions
von: Lee, Sangmin, et al.
Veröffentlicht: (2024) -
SocialGesture: Delving into Multi-person Gesture Understanding
von: Cao, Xu, et al.
Veröffentlicht: (2025) -
Leveraging Object Priors for Point Tracking
von: Boote, Bikram, et al.
Veröffentlicht: (2024) -
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
von: Lai, Bolin, et al.
Veröffentlicht: (2025)