Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Lei, Tito, Rubèn, Valveny, Ernest, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
Privacy-Aware Document Visual Question Answering
von: Tito, Rubèn, et al.
Veröffentlicht: (2023)
von: Tito, Rubèn, et al.
Veröffentlicht: (2023)
Federated Document Visual Question Answering: A Pilot Study
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
Reading in the Dark: Low-light Scene Text Recognition
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
Machine Unlearning for Document Classification
von: Kang, Lei, et al.
Veröffentlicht: (2024)
von: Kang, Lei, et al.
Veröffentlicht: (2024)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
von: Kang, Lei, et al.
Veröffentlicht: (2025)
von: Kang, Lei, et al.
Veröffentlicht: (2025)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
Multimodal Transformer for Comics Text-Cloze
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
Image-text matching for large-scale book collections
von: Llabrés, Artemis, et al.
Veröffentlicht: (2024)
von: Llabrés, Artemis, et al.
Veröffentlicht: (2024)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models
von: Kang, Lei, et al.
Veröffentlicht: (2024)
von: Kang, Lei, et al.
Veröffentlicht: (2024)
ComicsPAP: understanding comic strips by picking the correct panel
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
von: Gomez, Lluis, et al.
Veröffentlicht: (2014)
von: Gomez, Lluis, et al.
Veröffentlicht: (2014)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
von: Ortega, Marc Serra, et al.
Veröffentlicht: (2025)
von: Ortega, Marc Serra, et al.
Veröffentlicht: (2025)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
von: Mishra, Pritam, et al.
Veröffentlicht: (2025)
von: Mishra, Pritam, et al.
Veröffentlicht: (2025)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
Enhancing Document VQA Models via Retrieval-Augmented Generation
von: López, Eric, et al.
Veröffentlicht: (2025)
von: López, Eric, et al.
Veröffentlicht: (2025)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
von: Mishra, Pritam, et al.
Veröffentlicht: (2026)
von: Mishra, Pritam, et al.
Veröffentlicht: (2026)
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
von: Wang, Zining, et al.
Veröffentlicht: (2025)
von: Wang, Zining, et al.
Veröffentlicht: (2025)
Multimodal Integration of Human-Like Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
von: Indrehus, Kjetil, et al.
Veröffentlicht: (2026)
von: Indrehus, Kjetil, et al.
Veröffentlicht: (2026)
Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
von: Li, Zhifei, et al.
Veröffentlicht: (2025)
von: Li, Zhifei, et al.
Veröffentlicht: (2025)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
von: Lassoued, Aymen, et al.
Veröffentlicht: (2026)
von: Lassoued, Aymen, et al.
Veröffentlicht: (2026)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2024)
von: Tosato, Lucrezia, et al.
Veröffentlicht: (2024)
Reading Between the Lanes: Text VideoQA on the Road
von: Tom, George, et al.
Veröffentlicht: (2023)
von: Tom, George, et al.
Veröffentlicht: (2023)
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
von: Lin, Yuxin, et al.
Veröffentlicht: (2025)
von: Lin, Yuxin, et al.
Veröffentlicht: (2025)
ConFoThinking: Consolidated Focused Attention Driven Thinking for Visual Question Answering
von: Wu, Zhaodong, et al.
Veröffentlicht: (2026)
von: Wu, Zhaodong, et al.
Veröffentlicht: (2026)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
VoQA: Visual-only Question Answering
von: An, Jianing, et al.
Veröffentlicht: (2025)
von: An, Jianing, et al.
Veröffentlicht: (2025)
Targeted Visual Prompting for Medical Visual Question Answering
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
Joint Extraction Matters: Prompt-Based Visual Question Answering for Multi-Field Document Information Extraction
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
von: Loem, Mengsay, et al.
Veröffentlicht: (2025)
Visually Interpretable Subtask Reasoning for Visual Question Answering
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
von: Kim, Hongyeob, et al.
Veröffentlicht: (2025)
von: Kim, Hongyeob, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
von: Li, Zongmin, et al.
Veröffentlicht: (2026) -
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025) -
Privacy-Aware Document Visual Question Answering
von: Tito, Rubèn, et al.
Veröffentlicht: (2023) -
Federated Document Visual Question Answering: A Pilot Study
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024) -
Reading in the Dark: Low-light Scene Text Recognition
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)