Social Caption: Evaluating Social Understanding in Multimodal Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Thumu, Bhaavanaa, Mathur, Leena, Kebe, Youssouf, Morency, Louis-Philippe |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
par: Mathur, Leena, et autres
Publié: (2025)
par: Mathur, Leena, et autres
Publié: (2025)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
par: Mathur, Leena, et autres
Publié: (2024)
par: Mathur, Leena, et autres
Publié: (2024)
HEMM: Holistic Evaluation of Multimodal Foundation Models
par: Liang, Paul Pu, et autres
Publié: (2024)
par: Liang, Paul Pu, et autres
Publié: (2024)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
par: Zhou, Xuhui, et autres
Publié: (2023)
par: Zhou, Xuhui, et autres
Publié: (2023)
Omitted Variable Bias in Language Models Under Distribution Shift
par: Lin, Victoria, et autres
Publié: (2026)
par: Lin, Victoria, et autres
Publié: (2026)
Optimizing Language Models for Human Preferences is a Causal Inference Problem
par: Lin, Victoria, et autres
Publié: (2024)
par: Lin, Victoria, et autres
Publié: (2024)
LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?
par: Kebe, Gaoussou Youssouf, et autres
Publié: (2025)
par: Kebe, Gaoussou Youssouf, et autres
Publié: (2025)
American Sign Language Video to Text Translation
par: Roy, Parsheeta, et autres
Publié: (2024)
par: Roy, Parsheeta, et autres
Publié: (2024)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
par: Lee, Dong Won, et autres
Publié: (2024)
par: Lee, Dong Won, et autres
Publié: (2024)
IoT-LM: Large Multisensory Language Models for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2023)
par: Mo, Shentong, et autres
Publié: (2023)
Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
par: Liang, Paul Pu, et autres
Publié: (2023)
par: Liang, Paul Pu, et autres
Publié: (2023)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
par: Yanuka, Moran, et autres
Publié: (2024)
par: Yanuka, Moran, et autres
Publié: (2024)
Time Series Language Model for Descriptive Caption Generation
par: Trabelsi, Mohamed, et autres
Publié: (2025)
par: Trabelsi, Mohamed, et autres
Publié: (2025)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
par: Elchafei, Passant, et autres
Publié: (2025)
par: Elchafei, Passant, et autres
Publié: (2025)
NodeSynth: Socially Aligned Synthetic Data for AI Evaluation
par: Rashid, Qazi Mamunur, et autres
Publié: (2026)
par: Rashid, Qazi Mamunur, et autres
Publié: (2026)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
par: Lee, Nahyun, et autres
Publié: (2026)
par: Lee, Nahyun, et autres
Publié: (2026)
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
par: Munikoti, Sai, et autres
Publié: (2026)
par: Munikoti, Sai, et autres
Publié: (2026)
Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
par: Lee, Dong Won, et autres
Publié: (2025)
par: Lee, Dong Won, et autres
Publié: (2025)
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
par: Lee, Sangmin, et autres
Publié: (2024)
par: Lee, Sangmin, et autres
Publié: (2024)
Graphically Speaking: Unmasking Abuse in Social Media with Conversation Insights
par: Nouri, Célia, et autres
Publié: (2025)
par: Nouri, Célia, et autres
Publié: (2025)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
par: Yu, Haofei, et autres
Publié: (2023)
par: Yu, Haofei, et autres
Publié: (2023)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
par: Ma, Yunsheng, et autres
Publié: (2024)
par: Ma, Yunsheng, et autres
Publié: (2024)
Evaluating LLM Story Generation through Large-scale Network Analysis of Social Structures
par: Nonaka, Hiroshi, et autres
Publié: (2025)
par: Nonaka, Hiroshi, et autres
Publié: (2025)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
par: Miyai, Atsuyuki, et autres
Publié: (2024)
par: Miyai, Atsuyuki, et autres
Publié: (2024)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
par: Hu, Jiewen, et autres
Publié: (2025)
par: Hu, Jiewen, et autres
Publié: (2025)
Can Large Language Models Transform Computational Social Science?
par: Ziems, Caleb, et autres
Publié: (2023)
par: Ziems, Caleb, et autres
Publié: (2023)
Social Learning: Towards Collaborative Learning with Large Language Models
par: Mohtashami, Amirkeivan, et autres
Publié: (2023)
par: Mohtashami, Amirkeivan, et autres
Publié: (2023)
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
par: Dang, Vy Tuong, et autres
Publié: (2025)
par: Dang, Vy Tuong, et autres
Publié: (2025)
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data
par: Abdelkadir, Nuredin Ali, et autres
Publié: (2026)
par: Abdelkadir, Nuredin Ali, et autres
Publié: (2026)
Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media
par: Hasan, Khalid, et autres
Publié: (2025)
par: Hasan, Khalid, et autres
Publié: (2025)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
par: Singh, Ashish, et autres
Publié: (2023)
par: Singh, Ashish, et autres
Publié: (2023)
A Systematic Analysis on the Temporal Generalization of Language Models in Social Media
par: Ushio, Asahi, et autres
Publié: (2024)
par: Ushio, Asahi, et autres
Publié: (2024)
Multi-Stage Training for Abusive Comment Detection in Indic Languages
par: Rastogi, Pranshu, et autres
Publié: (2026)
par: Rastogi, Pranshu, et autres
Publié: (2026)
A Unified Understanding and Evaluation of Steering Methods
par: Im, Shawn, et autres
Publié: (2025)
par: Im, Shawn, et autres
Publié: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
par: Majumdar, Ayan, et autres
Publié: (2025)
par: Majumdar, Ayan, et autres
Publié: (2025)
Linear Alignment of Vision-language Models for Image Captioning
par: Paischer, Fabian, et autres
Publié: (2023)
par: Paischer, Fabian, et autres
Publié: (2023)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
par: Liu, Zhihang, et autres
Publié: (2025)
par: Liu, Zhihang, et autres
Publié: (2025)
Chem4DLLM: 4D Multimodal LLMs for Chemical Dynamics Understanding
par: Li, Xinyu, et autres
Publié: (2026)
par: Li, Xinyu, et autres
Publié: (2026)
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding
par: Cheshmi, Seyyed Saeid, et autres
Publié: (2026)
par: Cheshmi, Seyyed Saeid, et autres
Publié: (2026)
Documents similaires
-
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
par: Mathur, Leena, et autres
Publié: (2025) -
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
par: Mathur, Leena, et autres
Publié: (2024) -
HEMM: Holistic Evaluation of Multimodal Foundation Models
par: Liang, Paul Pu, et autres
Publié: (2024) -
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
par: Zhou, Xuhui, et autres
Publié: (2023) -
Omitted Variable Bias in Language Models Under Distribution Shift
par: Lin, Victoria, et autres
Publié: (2026)