Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
Fuente:
arXiv
Guardado en:
| Autores principales: | Haydarov, Kilichbek, Shen, Xiaoqian, Madasu, Avinash, Salem, Mahmoud, Li, Li-Jia, Elsayed, Gamaleldin, Elhoseiny, Mohamed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
StoryGPT-V: Large Language Models as Consistent Story Visualizers
por: Shen, Xiaoqian, et al.
Publicado: (2023)
por: Shen, Xiaoqian, et al.
Publicado: (2023)
A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity
por: Radwan, Yousef A., et al.
Publicado: (2026)
por: Radwan, Yousef A., et al.
Publicado: (2026)
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
por: Mohamed, Youssef, et al.
Publicado: (2024)
por: Mohamed, Youssef, et al.
Publicado: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
por: Shen, Xiaoqian, et al.
Publicado: (2025)
por: Shen, Xiaoqian, et al.
Publicado: (2025)
VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation
por: ZHU, Linan, et al.
Publicado: (2026)
por: ZHU, Linan, et al.
Publicado: (2026)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
por: Ahmed, Mahmoud, et al.
Publicado: (2025)
por: Ahmed, Mahmoud, et al.
Publicado: (2025)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
por: Madasu, Avinash, et al.
Publicado: (2023)
por: Madasu, Avinash, et al.
Publicado: (2023)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
por: Chowdhury, Sanjoy, et al.
Publicado: (2024)
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
por: Ye, Zilyu, et al.
Publicado: (2024)
por: Ye, Zilyu, et al.
Publicado: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
por: Yuan, Haobo, et al.
Publicado: (2025)
por: Yuan, Haobo, et al.
Publicado: (2025)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
Affective Flow Language Model for Emotional Support Conversation
por: Zou, Chenghui, et al.
Publicado: (2026)
por: Zou, Chenghui, et al.
Publicado: (2026)
RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
por: Li, Linfei, et al.
Publicado: (2026)
por: Li, Linfei, et al.
Publicado: (2026)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
por: Madasu, Avinash, et al.
Publicado: (2025)
por: Madasu, Avinash, et al.
Publicado: (2025)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
por: Madasu, Avinash, et al.
Publicado: (2025)
por: Madasu, Avinash, et al.
Publicado: (2025)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Affective Visualization Design: Leveraging the Emotional Impact of Data
por: Lan, Xingyu, et al.
Publicado: (2023)
por: Lan, Xingyu, et al.
Publicado: (2023)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
por: Li, Haocheng, et al.
Publicado: (2026)
por: Li, Haocheng, et al.
Publicado: (2026)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
por: Yu, Sungduk, et al.
Publicado: (2025)
por: Yu, Sungduk, et al.
Publicado: (2025)
RVTBench: A Benchmark for Visual Reasoning Tasks
por: Shen, Yiqing, et al.
Publicado: (2025)
por: Shen, Yiqing, et al.
Publicado: (2025)
Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
por: Satar, Burak, et al.
Publicado: (2025)
por: Satar, Burak, et al.
Publicado: (2025)
Affective Color Scales for Colormap Data Visualizations
por: Braun, Halle C., et al.
Publicado: (2025)
por: Braun, Halle C., et al.
Publicado: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
por: Lu, Lidong, et al.
Publicado: (2025)
por: Lu, Lidong, et al.
Publicado: (2025)
VGR: Visual Grounded Reasoning
por: Wang, Jiacong, et al.
Publicado: (2025)
por: Wang, Jiacong, et al.
Publicado: (2025)
Adaptive Masking Enhances Visual Grounding
por: Jia, Sen, et al.
Publicado: (2024)
por: Jia, Sen, et al.
Publicado: (2024)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
por: Park, Se Jin, et al.
Publicado: (2024)
por: Park, Se Jin, et al.
Publicado: (2024)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
por: Pang, Wei, et al.
Publicado: (2024)
por: Pang, Wei, et al.
Publicado: (2024)
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Learning from Reasoning Failures via Synthetic Data Generation
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2025)
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2025)
Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs
por: Mishra, Sandeep, et al.
Publicado: (2026)
por: Mishra, Sandeep, et al.
Publicado: (2026)
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
por: Guo, Guangfu, et al.
Publicado: (2026)
por: Guo, Guangfu, et al.
Publicado: (2026)
Quantifying and Enabling the Interpretability of CLIP-like Models
por: Madasu, Avinash, et al.
Publicado: (2024)
por: Madasu, Avinash, et al.
Publicado: (2024)
KOREYS TILIDA O'LCHOV BIRLIKLARI NOMLARINING TARKIBIY TUZILISHI VA HOSIL BO'LISH XUSUSIYATLARI
por: Haydarov, Jasur
Publicado: (2025)
por: Haydarov, Jasur
Publicado: (2025)
BOSHQARUV TIZIMINI TAKOMILLASHTIRISHDA AXBOROT-KOMMUNIKATSION TEXNOLOGIYALARNING O'RNI
por: Husanboy Haydarov
Publicado: (2025)
por: Husanboy Haydarov
Publicado: (2025)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
por: Felemban, Abdulwahab, et al.
Publicado: (2024)
por: Felemban, Abdulwahab, et al.
Publicado: (2024)
Neurotoxicidad en neonatos con hiperbilirrubinemia severa. Análisis de los factores de riesgo para neurotoxicidad en neonatos con ictericia severa
por: Rasha Gamaleldin
Publicado: (2012)
por: Rasha Gamaleldin
Publicado: (2012)
Ejemplares similares
-
StoryGPT-V: Large Language Models as Consistent Story Visualizers
por: Shen, Xiaoqian, et al.
Publicado: (2023) -
A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity
por: Radwan, Yousef A., et al.
Publicado: (2026) -
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
por: Mohamed, Youssef, et al.
Publicado: (2024) -
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
por: Abdelrahman, Eslam, et al.
Publicado: (2023) -
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
por: Shen, Xiaoqian, et al.
Publicado: (2025)