Gems: Group Emotion Profiling Through Multimodal Situational Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Kataria, Anubhav, Madan, Surbhi, Ghosh, Shreya, Gedeon, Tom, Dhall, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CSGaze: Context-aware Social Gaze Prediction
by: Madan, Surbhi, et al.
Published: (2025)
by: Madan, Surbhi, et al.
Published: (2025)
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding
by: Madan, Surbhi, et al.
Published: (2024)
by: Madan, Surbhi, et al.
Published: (2024)
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
by: Gupta, Parul, et al.
Published: (2025)
by: Gupta, Parul, et al.
Published: (2025)
MRAC Track 1: 2nd Workshop on Multimodal, Generative and Responsible Affective Computing
by: Ghosh, Shreya, et al.
Published: (2024)
by: Ghosh, Shreya, et al.
Published: (2024)
A Survey of Deep Learning for Group-level Emotion Recognition
by: Huang, Xiaohua, et al.
Published: (2024)
by: Huang, Xiaohua, et al.
Published: (2024)
Pavlok-Nudge: A Feedback Mechanism for Atomic Behaviour Modification with Snoring Usecase
by: Hasan, Md Rakibul, et al.
Published: (2023)
by: Hasan, Md Rakibul, et al.
Published: (2023)
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
by: Cai, Zhixi, et al.
Published: (2023)
by: Cai, Zhixi, et al.
Published: (2023)
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
What Does Softmax Probability Tell Us about Classifiers Ranking Across Diverse Test Conditions?
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Taylor Videos for Action Recognition
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Measuring Style Similarity in Diffusion Models
by: Somepalli, Gowthami, et al.
Published: (2024)
by: Somepalli, Gowthami, et al.
Published: (2024)
TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps
by: Raj, Arjun, et al.
Published: (2024)
by: Raj, Arjun, et al.
Published: (2024)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
An Enhanced Large Language Model For Cross Modal Query Understanding System Using DL-KeyBERT Based CAZSSCL-MPGPT
by: Singh, Shreya
Published: (2025)
by: Singh, Shreya
Published: (2025)
An Empirical Study Into What Matters for Calibrating Vision-Language Models
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
When Spatial meets Temporal in Action Recognition
by: Chen, Huilin, et al.
Published: (2024)
by: Chen, Huilin, et al.
Published: (2024)
SFANet: Spatial-Frequency Attention Network for Deepfake Detection
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025)
by: Agarwal, Vatsal, et al.
Published: (2025)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
Adaptive Multi-head Contrastive Learning
by: Wang, Lei, et al.
Published: (2023)
by: Wang, Lei, et al.
Published: (2023)
SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
by: Yee, Phyo Thet, et al.
Published: (2025)
by: Yee, Phyo Thet, et al.
Published: (2025)
Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition
by: Wang, Zeheng, et al.
Published: (2026)
by: Wang, Zeheng, et al.
Published: (2026)
7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition
by: Kollias, Dimitrios, et al.
Published: (2024)
by: Kollias, Dimitrios, et al.
Published: (2024)
1M-Deepfakes Detection Challenge
by: Cai, Zhixi, et al.
Published: (2024)
by: Cai, Zhixi, et al.
Published: (2024)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal-Viewpoint Alignment
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
by: Ding, Dexuan, et al.
Published: (2024)
by: Ding, Dexuan, et al.
Published: (2024)
How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?
by: Purohit, Mirali, et al.
Published: (2025)
by: Purohit, Mirali, et al.
Published: (2025)
iWatchRoadv2: Pothole Detection, Geospatial Mapping, and Intelligent Road Governance
by: Sahoo, Rishi Raj, et al.
Published: (2025)
by: Sahoo, Rishi Raj, et al.
Published: (2025)
iWatchRoad: Scalable Detection and Geospatial Visualization of Potholes for Smart Cities
by: Sahoo, Rishi Raj, et al.
Published: (2025)
by: Sahoo, Rishi Raj, et al.
Published: (2025)
LayLens: Improving Deepfake Understanding through Simplified Explanations
by: Narang, Abhijeet, et al.
Published: (2025)
by: Narang, Abhijeet, et al.
Published: (2025)
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
Bayesian Multi-Scale Neural Network for Crowd Counting
by: Sagar, Abhinav
Published: (2020)
by: Sagar, Abhinav
Published: (2020)
Privacy-Preserving Empathy Detection in Video Interactions
by: Hasan, Md Rakibul, et al.
Published: (2025)
by: Hasan, Md Rakibul, et al.
Published: (2025)
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
by: Kaur, Amandeep, et al.
Published: (2026)
by: Kaur, Amandeep, et al.
Published: (2026)
Emolysis: A Multimodal Open-Source Group Emotion Analysis and Visualization Toolkit
by: Ghosh, Shreya, et al.
Published: (2023)
by: Ghosh, Shreya, et al.
Published: (2023)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
by: Naghashyar, Lachin, et al.
Published: (2026)
by: Naghashyar, Lachin, et al.
Published: (2026)
Similar Items
-
CSGaze: Context-aware Social Gaze Prediction
by: Madan, Surbhi, et al.
Published: (2025) -
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding
by: Madan, Surbhi, et al.
Published: (2024) -
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
by: Gupta, Parul, et al.
Published: (2025) -
MRAC Track 1: 2nd Workshop on Multimodal, Generative and Responsible Affective Computing
by: Ghosh, Shreya, et al.
Published: (2024) -
A Survey of Deep Learning for Group-level Emotion Recognition
by: Huang, Xiaohua, et al.
Published: (2024)