MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Madan, Surbhi, Ghosh, Shreya, Sookha, Lownish Rai, Ganaie, M. A., Subramanian, Ramanathan, Dhall, Abhinav, Gedeon, Tom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CSGaze: Context-aware Social Gaze Prediction
von: Madan, Surbhi, et al.
Veröffentlicht: (2025)
von: Madan, Surbhi, et al.
Veröffentlicht: (2025)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques
von: Sookha, Lownish Rai, et al.
Veröffentlicht: (2025)
von: Sookha, Lownish Rai, et al.
Veröffentlicht: (2025)
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
von: Gupta, Parul, et al.
Veröffentlicht: (2025)
von: Gupta, Parul, et al.
Veröffentlicht: (2025)
Gems: Group Emotion Profiling Through Multimodal Situational Understanding
von: Kataria, Anubhav, et al.
Veröffentlicht: (2025)
von: Kataria, Anubhav, et al.
Veröffentlicht: (2025)
SFANet: Spatial-Frequency Attention Network for Deepfake Detection
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
LayLens: Improving Deepfake Understanding through Simplified Explanations
von: Narang, Abhijeet, et al.
Veröffentlicht: (2025)
von: Narang, Abhijeet, et al.
Veröffentlicht: (2025)
EEG-based Cognitive Load Estimation of Acoustic Parameters for Data Sonification
von: Sharma, Gulshan, et al.
Veröffentlicht: (2024)
von: Sharma, Gulshan, et al.
Veröffentlicht: (2024)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
von: Gu, Yimeng, et al.
Veröffentlicht: (2025)
von: Gu, Yimeng, et al.
Veröffentlicht: (2025)
QuMATL: Query-based Multi-annotator Tendency Learning
von: Zhang, Liyun, et al.
Veröffentlicht: (2025)
von: Zhang, Liyun, et al.
Veröffentlicht: (2025)
Emolysis: A Multimodal Open-Source Group Emotion Analysis and Visualization Toolkit
von: Ghosh, Shreya, et al.
Veröffentlicht: (2023)
von: Ghosh, Shreya, et al.
Veröffentlicht: (2023)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
MRAC Track 1: 2nd Workshop on Multimodal, Generative and Responsible Affective Computing
von: Ghosh, Shreya, et al.
Veröffentlicht: (2024)
von: Ghosh, Shreya, et al.
Veröffentlicht: (2024)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
von: Li, Fanxiao, et al.
Veröffentlicht: (2025)
von: Li, Fanxiao, et al.
Veröffentlicht: (2025)
Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
SimLabel: Similarity-Weighted Iterative Framework for Multi-annotator Learning with Missing Annotations
von: Zhang, Liyun, et al.
Veröffentlicht: (2025)
von: Zhang, Liyun, et al.
Veröffentlicht: (2025)
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
Pavlok-Nudge: A Feedback Mechanism for Atomic Behaviour Modification with Snoring Usecase
von: Hasan, Md Rakibul, et al.
Veröffentlicht: (2023)
von: Hasan, Md Rakibul, et al.
Veröffentlicht: (2023)
Semantically Consistent Person Image Generation
von: Roy, Prasun, et al.
Veröffentlicht: (2023)
von: Roy, Prasun, et al.
Veröffentlicht: (2023)
Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
von: Ghosh, Indrajeet, et al.
Veröffentlicht: (2025)
von: Ghosh, Indrajeet, et al.
Veröffentlicht: (2025)
Prototypical Prompting for Text-to-image Person Re-identification
von: Yan, Shuanglin, et al.
Veröffentlicht: (2024)
von: Yan, Shuanglin, et al.
Veröffentlicht: (2024)
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
von: Ancarani, Elisa, et al.
Veröffentlicht: (2025)
von: Ancarani, Elisa, et al.
Veröffentlicht: (2025)
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
Scene Aware Person Image Generation through Global Contextual Conditioning
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
The Future is Meta: Metadata, Formats and Perspectives towards Interactive and Personalized AV Content
von: Weller, Alexander, et al.
Veröffentlicht: (2024)
von: Weller, Alexander, et al.
Veröffentlicht: (2024)
EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
von: Girmaji, Rohit, et al.
Veröffentlicht: (2025)
von: Girmaji, Rohit, et al.
Veröffentlicht: (2025)
Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
von: Yang, Jasmine, et al.
Veröffentlicht: (2026)
von: Yang, Jasmine, et al.
Veröffentlicht: (2026)
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
von: Zhang, Beibei, et al.
Veröffentlicht: (2025)
von: Zhang, Beibei, et al.
Veröffentlicht: (2025)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
von: Deng, Weihui, et al.
Veröffentlicht: (2024)
von: Deng, Weihui, et al.
Veröffentlicht: (2024)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025)
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025)
How to Cache Important Contents for Multi-modal Service in Dynamic Networks: A DRL-based Caching Scheme
von: Zhang, Zhe, et al.
Veröffentlicht: (2024)
von: Zhang, Zhe, et al.
Veröffentlicht: (2024)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
von: Bai, Hayes, et al.
Veröffentlicht: (2026)
von: Bai, Hayes, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CSGaze: Context-aware Social Gaze Prediction
von: Madan, Surbhi, et al.
Veröffentlicht: (2025) -
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025) -
A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques
von: Sookha, Lownish Rai, et al.
Veröffentlicht: (2025) -
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
von: Gupta, Parul, et al.
Veröffentlicht: (2025) -
Gems: Group Emotion Profiling Through Multimodal Situational Understanding
von: Kataria, Anubhav, et al.
Veröffentlicht: (2025)