BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mao, Yuanyuan, Lin, Xin, Ni, Qin, He, Liang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
Memory-Centric Embodied Question Answering
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
VidCtx: Context-aware Video Question Answering with Image Models
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
von: Tian, Chong, et al.
Veröffentlicht: (2026)
von: Tian, Chong, et al.
Veröffentlicht: (2026)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models
von: Lin, Yuxiang, et al.
Veröffentlicht: (2025)
von: Lin, Yuxiang, et al.
Veröffentlicht: (2025)
Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts
von: Cole, Adam, et al.
Veröffentlicht: (2025)
von: Cole, Adam, et al.
Veröffentlicht: (2025)
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline
von: Oh, Minwoo, et al.
Veröffentlicht: (2025)
von: Oh, Minwoo, et al.
Veröffentlicht: (2025)
Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
von: Han, Wei, et al.
Veröffentlicht: (2023)
von: Han, Wei, et al.
Veröffentlicht: (2023)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Plasticity-Aware Mixture of Experts for Learning Under QoE Shifts in Adaptive Video Streaming
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
MindFuse: Towards GenAI Explainability in Marketing Strategy Co-Creation
von: Farseev, Aleksandr, et al.
Veröffentlicht: (2025)
von: Farseev, Aleksandr, et al.
Veröffentlicht: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
von: Yang, Zhe, et al.
Veröffentlicht: (2024)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
von: Wang, Yihao, et al.
Veröffentlicht: (2024)
von: Wang, Yihao, et al.
Veröffentlicht: (2024)
Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey
von: Lin, Qika, et al.
Veröffentlicht: (2024)
von: Lin, Qika, et al.
Veröffentlicht: (2024)
HiQuE: Hierarchical Question Embedding Network for Multimodal Depression Detection
von: Jung, Juho, et al.
Veröffentlicht: (2024)
von: Jung, Juho, et al.
Veröffentlicht: (2024)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
Semantic-Guided Unsupervised Video Summarization
von: Liu, Haizhou, et al.
Veröffentlicht: (2026)
von: Liu, Haizhou, et al.
Veröffentlicht: (2026)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
A New Dataset and Benchmark for Grounding Multimodal Misinformation
von: Yang, Bingjian, et al.
Veröffentlicht: (2025)
von: Yang, Bingjian, et al.
Veröffentlicht: (2025)
Towards Open-Vocabulary Video Semantic Segmentation
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
End-to-End Learning-based Video Streaming Enhancement Pipeline: A Generative AI Approach
von: Artioli, Emanuele, et al.
Veröffentlicht: (2025)
von: Artioli, Emanuele, et al.
Veröffentlicht: (2025)
MM-HSD: Multi-Modal Hate Speech Detection in Videos
von: Céspedes-Sarrias, Berta, et al.
Veröffentlicht: (2025)
von: Céspedes-Sarrias, Berta, et al.
Veröffentlicht: (2025)
FineFake: A Knowledge-Enriched Dataset for Fine-Grained Multi-Domain Fake News Detection
von: Zhou, Ziyi, et al.
Veröffentlicht: (2024)
von: Zhou, Ziyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024) -
Memory-Centric Embodied Question Answering
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025) -
VidCtx: Context-aware Video Question Answering with Image Models
von: Goulas, Andreas, et al.
Veröffentlicht: (2024) -
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023) -
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)