Enhanced Textual Feature Extraction for Visual Question Answering: A Simple Convolutional Approach
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zhilin, Wu, Fangyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
por: Thai, Triet Minh, et al.
Publicado: (2023)
por: Thai, Triet Minh, et al.
Publicado: (2023)
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
por: Nikandrou, Malvina, et al.
Publicado: (2024)
por: Nikandrou, Malvina, et al.
Publicado: (2024)
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
por: Nguyen, Ngoc Son, et al.
Publicado: (2024)
por: Nguyen, Ngoc Son, et al.
Publicado: (2024)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
por: Zhang, Xiaoman, et al.
Publicado: (2023)
por: Zhang, Xiaoman, et al.
Publicado: (2023)
A Simple LLM Framework for Long-Range Video Question-Answering
por: Zhang, Ce, et al.
Publicado: (2023)
por: Zhang, Ce, et al.
Publicado: (2023)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
por: Xue, Junxiao, et al.
Publicado: (2024)
por: Xue, Junxiao, et al.
Publicado: (2024)
Efficient Bilinear Attention-based Fusion for Medical Visual Question Answering
por: Zhang, Zhilin, et al.
Publicado: (2024)
por: Zhang, Zhilin, et al.
Publicado: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
por: Kim, Hongyeob, et al.
Publicado: (2025)
por: Kim, Hongyeob, et al.
Publicado: (2025)
Joint Extraction Matters: Prompt-Based Visual Question Answering for Multi-Field Document Information Extraction
por: Loem, Mengsay, et al.
Publicado: (2025)
por: Loem, Mengsay, et al.
Publicado: (2025)
Harmonizing Feature Maps: A Graph Convolutional Approach for Enhancing Adversarial Robustness
por: Zhang, Kejia, et al.
Publicado: (2024)
por: Zhang, Kejia, et al.
Publicado: (2024)
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
por: Tian, Yuanhe, et al.
Publicado: (2025)
por: Tian, Yuanhe, et al.
Publicado: (2025)
Selectively Answering Visual Questions
por: Eisenschlos, Julian Martin, et al.
Publicado: (2024)
por: Eisenschlos, Julian Martin, et al.
Publicado: (2024)
Targeted Visual Prompting for Medical Visual Question Answering
por: Tascon-Morales, Sergio, et al.
Publicado: (2024)
por: Tascon-Morales, Sergio, et al.
Publicado: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
por: Ishmam, Md Farhan, et al.
Publicado: (2024)
por: Ishmam, Md Farhan, et al.
Publicado: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
por: Cheng, Yu, et al.
Publicado: (2025)
por: Cheng, Yu, et al.
Publicado: (2025)
Questioning the Stability of Visual Question Answering
por: Rosenfeld, Amir, et al.
Publicado: (2025)
por: Rosenfeld, Amir, et al.
Publicado: (2025)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
por: Xu, Quanxing, et al.
Publicado: (2026)
por: Xu, Quanxing, et al.
Publicado: (2026)
Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering
por: Tao, Qian, et al.
Publicado: (2025)
por: Tao, Qian, et al.
Publicado: (2025)
Geospatial Chain of Thought Reasoning for Enhanced Visual Question Answering on Satellite Imagery
por: Shanker, Shambhavi, et al.
Publicado: (2025)
por: Shanker, Shambhavi, et al.
Publicado: (2025)
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
por: Movva, Prahitha, et al.
Publicado: (2025)
por: Movva, Prahitha, et al.
Publicado: (2025)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
por: Wang, Yanling, et al.
Publicado: (2025)
por: Wang, Yanling, et al.
Publicado: (2025)
Object Retrieval for Visual Question Answering with Outside Knowledge
por: Kan, Shichao, et al.
Publicado: (2024)
por: Kan, Shichao, et al.
Publicado: (2024)
Hallucination Benchmark in Medical Visual Question Answering
por: Wu, Jinge, et al.
Publicado: (2024)
por: Wu, Jinge, et al.
Publicado: (2024)
Multimodal Rationales for Explainable Visual Question Answering
por: Li, Kun, et al.
Publicado: (2024)
por: Li, Kun, et al.
Publicado: (2024)
Evaluating Variance in Visual Question Answering Benchmarks
por: SR, Nikitha
Publicado: (2025)
por: SR, Nikitha
Publicado: (2025)
DriveLM: Driving with Graph Visual Question Answering
por: Sima, Chonghao, et al.
Publicado: (2023)
por: Sima, Chonghao, et al.
Publicado: (2023)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
por: Özdemir, Övgü, et al.
Publicado: (2024)
por: Özdemir, Övgü, et al.
Publicado: (2024)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
por: Romero, David, et al.
Publicado: (2024)
por: Romero, David, et al.
Publicado: (2024)
Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
por: Wang, Zhifeng, et al.
Publicado: (2025)
por: Wang, Zhifeng, et al.
Publicado: (2025)
Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering
por: Wang, Zeqing, et al.
Publicado: (2023)
por: Wang, Zeqing, et al.
Publicado: (2023)
QIRL: Boosting Visual Question Answering via Optimized Question-Image Relation Learning
por: Xu, Quanxing, et al.
Publicado: (2025)
por: Xu, Quanxing, et al.
Publicado: (2025)
VoQA: Visual-only Question Answering
por: An, Jianing, et al.
Publicado: (2025)
por: An, Jianing, et al.
Publicado: (2025)
Reconstruction as a Bridge for Event-Based Visual Question Answering
por: Lou, Hanyue, et al.
Publicado: (2025)
por: Lou, Hanyue, et al.
Publicado: (2025)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
por: Zhang, Zhengxuan, et al.
Publicado: (2025)
por: Zhang, Zhengxuan, et al.
Publicado: (2025)
Saliency Guided Longitudinal Medical Visual Question Answering
por: Wu, Jialin, et al.
Publicado: (2025)
por: Wu, Jialin, et al.
Publicado: (2025)
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
por: Yang, Tianyu, et al.
Publicado: (2024)
por: Yang, Tianyu, et al.
Publicado: (2024)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
por: Kabir, Raihan, et al.
Publicado: (2024)
por: Kabir, Raihan, et al.
Publicado: (2024)
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
por: Zheng, Peiru, et al.
Publicado: (2024)
por: Zheng, Peiru, et al.
Publicado: (2024)
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
por: Wu, Xiangyu, et al.
Publicado: (2024)
por: Wu, Xiangyu, et al.
Publicado: (2024)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
por: Peng, Jingwei, et al.
Publicado: (2025)
por: Peng, Jingwei, et al.
Publicado: (2025)
Ejemplares similares
-
Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
por: Thai, Triet Minh, et al.
Publicado: (2023) -
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
por: Nikandrou, Malvina, et al.
Publicado: (2024) -
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
por: Nguyen, Ngoc Son, et al.
Publicado: (2024) -
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
por: Zhang, Xiaoman, et al.
Publicado: (2023) -
A Simple LLM Framework for Long-Range Video Question-Answering
por: Zhang, Ce, et al.
Publicado: (2023)