ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images
Fuente:
arXiv
Salvato in:
| Autori principali: | Van Nguyen, Quan, Tran, Dan Quang, Pham, Huy Quang, Nguyen, Thang Kien-Bao, Nguyen, Nghia Hieu, Van Nguyen, Kiet, Nguyen, Ngan Luu-Thuy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
ViMMRC 2.0 -- Enhancing Machine Reading Comprehension on Vietnamese Literature Text
di: Luu, Son T., et al.
Pubblicazione: (2023)
di: Luu, Son T., et al.
Pubblicazione: (2023)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
EVJVQA Challenge: Multilingual Visual Question Answering
di: Nguyen, Ngan Luu-Thuy, et al.
Pubblicazione: (2023)
di: Nguyen, Ngan Luu-Thuy, et al.
Pubblicazione: (2023)
Vietnamese AI Generated Text Detection
di: Tran, Quang-Dan, et al.
Pubblicazione: (2024)
di: Tran, Quang-Dan, et al.
Pubblicazione: (2024)
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
di: Tran, Hung Quang, et al.
Pubblicazione: (2026)
di: Tran, Hung Quang, et al.
Pubblicazione: (2026)
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
di: Hoang, Quan Ngoc, et al.
Pubblicazione: (2026)
di: Hoang, Quan Ngoc, et al.
Pubblicazione: (2026)
Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2026)
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2026)
Coreference Resolution for Vietnamese Narrative Texts
di: Tran, Hieu-Dai, et al.
Pubblicazione: (2025)
di: Tran, Hieu-Dai, et al.
Pubblicazione: (2025)
VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
di: Luu, Son T., et al.
Pubblicazione: (2025)
di: Luu, Son T., et al.
Pubblicazione: (2025)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
di: Nguyen, Hieu Minh, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu Minh, et al.
Pubblicazione: (2025)
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases
di: Nguyen, Sang Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Sang Quang, et al.
Pubblicazione: (2025)
A New Benchmark Dataset and Mixture-of-Experts Language Models for Adversarial Natural Language Inference in Vietnamese
di: Van Huynh, Tin, et al.
Pubblicazione: (2024)
di: Van Huynh, Tin, et al.
Pubblicazione: (2024)
ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition
di: Nguyen, Khoa Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Khoa Anh, et al.
Pubblicazione: (2026)
ViMultiChoice: Toward a Method That Gives Explanation for Multiple-Choice Reading Comprehension in Vietnamese
di: Cao, Trung Tien, et al.
Pubblicazione: (2026)
di: Cao, Trung Tien, et al.
Pubblicazione: (2026)
Optimizing Legal Document Retrieval in Vietnamese with Semi-Hard Negative Mining
di: Le, Van-Hoang, et al.
Pubblicazione: (2025)
di: Le, Van-Hoang, et al.
Pubblicazione: (2025)
ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks
di: Van Huynh, Tin, et al.
Pubblicazione: (2026)
di: Van Huynh, Tin, et al.
Pubblicazione: (2026)
An Attempt to Develop a Neural Parser based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese
di: Nguyen, Duc-Vu, et al.
Pubblicazione: (2024)
di: Nguyen, Duc-Vu, et al.
Pubblicazione: (2024)
Transformer-Based Contextualized Language Models Joint with Neural Networks for Natural Language Inference in Vietnamese
di: Nguyen, Dat Van-Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Dat Van-Thanh, et al.
Pubblicazione: (2024)
LiGT: Layout-infused Generative Transformer for Visual Question Answering on Vietnamese Receipts
di: Le, Thanh-Phong, et al.
Pubblicazione: (2025)
di: Le, Thanh-Phong, et al.
Pubblicazione: (2025)
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
di: Nguyen, Khoi Anh, et al.
Pubblicazione: (2025)
di: Nguyen, Khoi Anh, et al.
Pubblicazione: (2025)
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
di: Ngo, Thinh Phuoc, et al.
Pubblicazione: (2024)
di: Ngo, Thinh Phuoc, et al.
Pubblicazione: (2024)
UIT-DarkCow team at ImageCLEFmedical Caption 2024: Diagnostic Captioning for Radiology Images Efficiency with Transformer Models
di: Van Nguyen, Quan, et al.
Pubblicazione: (2024)
di: Van Nguyen, Quan, et al.
Pubblicazione: (2024)
Vietnamese Legal Information Retrieval in Question-Answering System
di: Ba, Thiem Nguyen, et al.
Pubblicazione: (2024)
di: Ba, Thiem Nguyen, et al.
Pubblicazione: (2024)
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
di: Nguyen, Thanh-Nhi, et al.
Pubblicazione: (2024)
di: Nguyen, Thanh-Nhi, et al.
Pubblicazione: (2024)
VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding
di: Do, Phong Nguyen-Thuan, et al.
Pubblicazione: (2024)
di: Do, Phong Nguyen-Thuan, et al.
Pubblicazione: (2024)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
di: Van-Dinh, Tue-Thu, et al.
Pubblicazione: (2025)
di: Van-Dinh, Tue-Thu, et al.
Pubblicazione: (2025)
ViFactCheck: A New Benchmark Dataset and Methods for Multi-domain News Fact-Checking in Vietnamese
di: Hoa, Tran Thai, et al.
Pubblicazione: (2024)
di: Hoa, Tran Thai, et al.
Pubblicazione: (2024)
ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation
di: Nguyen, Truc Mai-Thanh, et al.
Pubblicazione: (2025)
di: Nguyen, Truc Mai-Thanh, et al.
Pubblicazione: (2025)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
di: Nguyen, Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Hieu, et al.
Pubblicazione: (2024)
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2025)
Can LLMs Play Ô Ăn Quan Game? A Study of Multi-Step Planning and Decision Making
di: Nguyen, Sang Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Sang Quang, et al.
Pubblicazione: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
di: Tuong, Nguyen Anh, et al.
Pubblicazione: (2026)
di: Tuong, Nguyen Anh, et al.
Pubblicazione: (2026)
ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard Translation
di: Ta, Khoa Anh, et al.
Pubblicazione: (2026)
di: Ta, Khoa Anh, et al.
Pubblicazione: (2026)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models
di: Nguyen, Duy Vu Minh, et al.
Pubblicazione: (2026)
di: Nguyen, Duy Vu Minh, et al.
Pubblicazione: (2026)
New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis
di: Nguyen, Quy Hoang, et al.
Pubblicazione: (2024)
di: Nguyen, Quy Hoang, et al.
Pubblicazione: (2024)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
di: Pham, Anh-Cuong, et al.
Pubblicazione: (2024)
di: Pham, Anh-Cuong, et al.
Pubblicazione: (2024)
ViLegalNLI: Natural Language Inference for Vietnamese Legal Texts
di: Duong, Nhung Thi-Hong, et al.
Pubblicazione: (2026)
di: Duong, Nhung Thi-Hong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
di: Pham, Huy Quang, et al.
Pubblicazione: (2024) -
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024) -
ViMMRC 2.0 -- Enhancing Machine Reading Comprehension on Vietnamese Literature Text
di: Luu, Son T., et al.
Pubblicazione: (2023) -
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026) -
EVJVQA Challenge: Multilingual Visual Question Answering
di: Nguyen, Ngan Luu-Thuy, et al.
Pubblicazione: (2023)