Visual question answering: from early developments to recent advances -- a survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huynh, Ngoc Dung, Bouadjenek, Mohamed Reda, Aryal, Sunil, Razzak, Imran, Hacid, Hakim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2024)
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2024)
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
Margin-bounded Confidence Scores for Out-of-Distribution Detection
von: Tamang, Lakpa D., et al.
Veröffentlicht: (2024)
von: Tamang, Lakpa D., et al.
Veröffentlicht: (2024)
Omnidirectional Video Super-Resolution using Deep Learning
von: Baniya, Arbind Agrahari, et al.
Veröffentlicht: (2025)
von: Baniya, Arbind Agrahari, et al.
Veröffentlicht: (2025)
Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles
von: Belgoumri, Mohammed Djameleddine, et al.
Veröffentlicht: (2025)
von: Belgoumri, Mohammed Djameleddine, et al.
Veröffentlicht: (2025)
TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
Deep Learning for Sports Video Event Detection: Tasks, Datasets, Methods, and Challenges
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
von: Gong, Zixuan, et al.
Veröffentlicht: (2024)
von: Gong, Zixuan, et al.
Veröffentlicht: (2024)
Causal Debiasing for Visual Commonsense Reasoning
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
Deep learning for 3D human pose estimation and mesh recovery: A survey
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
Learning Brain Representation with Hierarchical Visual Embeddings
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
Extending Visual Dynamics for Video-to-Music Generation
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Language-Guided Diffusion Model for Visual Grounding
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
CLIP Brings Better Features to Visual Aesthetics Learners
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
von: Wan, Siqi, et al.
Veröffentlicht: (2025)
von: Wan, Siqi, et al.
Veröffentlicht: (2025)
Noise-Tolerant Learning for Audio-Visual Action Recognition
von: Han, Haochen, et al.
Veröffentlicht: (2022)
von: Han, Haochen, et al.
Veröffentlicht: (2022)
Holistic Visual-Textual Sentiment Analysis with Prior Models
von: Chen, Junyu, et al.
Veröffentlicht: (2022)
von: Chen, Junyu, et al.
Veröffentlicht: (2022)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
von: Mao, Qingyang, et al.
Veröffentlicht: (2025)
von: Mao, Qingyang, et al.
Veröffentlicht: (2025)
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
Towards Open-Vocabulary Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
Class Agnostic Instance-level Descriptor for Visual Instance Search
von: Sun, Qi-Ying, et al.
Veröffentlicht: (2025)
von: Sun, Qi-Ying, et al.
Veröffentlicht: (2025)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding
von: Bao, Guangyin, et al.
Veröffentlicht: (2024)
von: Bao, Guangyin, et al.
Veröffentlicht: (2024)
Scaling Audio-Visual Quality Assessment Dataset via Crowdsourcing
von: Yang, Renyu, et al.
Veröffentlicht: (2026)
von: Yang, Renyu, et al.
Veröffentlicht: (2026)
Learning from Silence and Noise for Visual Sound Source Localization
von: Juanola, Xavier, et al.
Veröffentlicht: (2025)
von: Juanola, Xavier, et al.
Veröffentlicht: (2025)
Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation
von: Chen, Yuanhong, et al.
Veröffentlicht: (2023)
von: Chen, Yuanhong, et al.
Veröffentlicht: (2023)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
von: Xu, Yue, et al.
Veröffentlicht: (2024)
von: Xu, Yue, et al.
Veröffentlicht: (2024)
StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
Patch-level Sounding Object Tracking for Audio-Visual Question Answering
von: Li, Zhangbin, et al.
Veröffentlicht: (2024)
von: Li, Zhangbin, et al.
Veröffentlicht: (2024)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering
von: Fan, Xinqi, et al.
Veröffentlicht: (2026)
von: Fan, Xinqi, et al.
Veröffentlicht: (2026)
QPT V2: Masked Image Modeling Advances Visual Scoring
von: Xie, Qizhi, et al.
Veröffentlicht: (2024)
von: Xie, Qizhi, et al.
Veröffentlicht: (2024)
Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
von: E, Shaojun, et al.
Veröffentlicht: (2025)
von: E, Shaojun, et al.
Veröffentlicht: (2025)
Q-Adapt: Adapting LMM for Visual Quality Assessment with Progressive Instruction Tuning
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2024) -
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025) -
Margin-bounded Confidence Scores for Out-of-Distribution Detection
von: Tamang, Lakpa D., et al.
Veröffentlicht: (2024) -
Omnidirectional Video Super-Resolution using Deep Learning
von: Baniya, Arbind Agrahari, et al.
Veröffentlicht: (2025) -
Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles
von: Belgoumri, Mohammed Djameleddine, et al.
Veröffentlicht: (2025)