ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Nghia Hieu, Quan, Tho Thanh, Nguyen, Ngan Luu-Thuy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
by: Pham, Huy Quang, et al.
Published: (2024)
by: Pham, Huy Quang, et al.
Published: (2024)
ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition
by: Nguyen, Khoa Anh, et al.
Published: (2026)
by: Nguyen, Khoa Anh, et al.
Published: (2026)
ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images
by: Van Nguyen, Quan, et al.
Published: (2024)
by: Van Nguyen, Quan, et al.
Published: (2024)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
Coreference Resolution for Vietnamese Narrative Texts
by: Tran, Hieu-Dai, et al.
Published: (2025)
by: Tran, Hieu-Dai, et al.
Published: (2025)
EVJVQA Challenge: Multilingual Visual Question Answering
by: Nguyen, Ngan Luu-Thuy, et al.
Published: (2023)
by: Nguyen, Ngan Luu-Thuy, et al.
Published: (2023)
LiGT: Layout-infused Generative Transformer for Visual Question Answering on Vietnamese Receipts
by: Le, Thanh-Phong, et al.
Published: (2025)
by: Le, Thanh-Phong, et al.
Published: (2025)
ViMultiChoice: Toward a Method That Gives Explanation for Multiple-Choice Reading Comprehension in Vietnamese
by: Cao, Trung Tien, et al.
Published: (2026)
by: Cao, Trung Tien, et al.
Published: (2026)
ViMMRC 2.0 -- Enhancing Machine Reading Comprehension on Vietnamese Literature Text
by: Luu, Son T., et al.
Published: (2023)
by: Luu, Son T., et al.
Published: (2023)
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
by: Hoang, Quan Ngoc, et al.
Published: (2026)
by: Hoang, Quan Ngoc, et al.
Published: (2026)
Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese
by: Nguyen, Nghia Hieu, et al.
Published: (2026)
by: Nguyen, Nghia Hieu, et al.
Published: (2026)
An Attempt to Develop a Neural Parser based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese
by: Nguyen, Duc-Vu, et al.
Published: (2024)
by: Nguyen, Duc-Vu, et al.
Published: (2024)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
by: Nguyen, Hieu, et al.
Published: (2024)
by: Nguyen, Hieu, et al.
Published: (2024)
ViHateT5: Enhancing Hate Speech Detection in Vietnamese With A Unified Text-to-Text Transformer Model
by: Nguyen, Luan Thanh
Published: (2024)
by: Nguyen, Luan Thanh
Published: (2024)
VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
by: Luu, Son T., et al.
Published: (2025)
by: Luu, Son T., et al.
Published: (2025)
Transformer-Based Contextualized Language Models Joint with Neural Networks for Natural Language Inference in Vietnamese
by: Nguyen, Dat Van-Thanh, et al.
Published: (2024)
by: Nguyen, Dat Van-Thanh, et al.
Published: (2024)
ViHERMES: A Graph-Grounded Multihop Question Answering Benchmark and System for Vietnamese Healthcare Regulations
by: Nguyen, Long S. T., et al.
Published: (2026)
by: Nguyen, Long S. T., et al.
Published: (2026)
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
by: Nguyen, Thanh-Nhi, et al.
Published: (2024)
by: Nguyen, Thanh-Nhi, et al.
Published: (2024)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
by: Nguyen, Hieu Minh, et al.
Published: (2025)
by: Nguyen, Hieu Minh, et al.
Published: (2025)
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
by: Nguyen, Ngoc Son, et al.
Published: (2024)
by: Nguyen, Ngoc Son, et al.
Published: (2024)
A New Benchmark Dataset and Mixture-of-Experts Language Models for Adversarial Natural Language Inference in Vietnamese
by: Van Huynh, Tin, et al.
Published: (2024)
by: Van Huynh, Tin, et al.
Published: (2024)
Enriching and Controlling Global Semantics for Text Summarization
by: Nguyen, Thong, et al.
Published: (2021)
by: Nguyen, Thong, et al.
Published: (2021)
Leveraging Sentence-oriented Augmentation and Transformer-Based Architecture for Vietnamese-Bahnaric Translation
by: Nguyen, Tan Sang, et al.
Published: (2026)
by: Nguyen, Tan Sang, et al.
Published: (2026)
Optimizing Legal Document Retrieval in Vietnamese with Semi-Hard Negative Mining
by: Le, Van-Hoang, et al.
Published: (2025)
by: Le, Van-Hoang, et al.
Published: (2025)
ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
by: Dang, Phuong-Nam, et al.
Published: (2025)
by: Dang, Phuong-Nam, et al.
Published: (2025)
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
by: Nguyen, Tan-Minh, et al.
Published: (2025)
by: Nguyen, Tan-Minh, et al.
Published: (2025)
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
by: Tran, Hung Quang, et al.
Published: (2026)
by: Tran, Hung Quang, et al.
Published: (2026)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
by: Tuong, Nguyen Anh, et al.
Published: (2026)
by: Tuong, Nguyen Anh, et al.
Published: (2026)
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
by: Ngo, Thinh Phuoc, et al.
Published: (2024)
ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks
by: Van Huynh, Tin, et al.
Published: (2026)
by: Van Huynh, Tin, et al.
Published: (2026)
ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation
by: Nguyen, Truc Mai-Thanh, et al.
Published: (2025)
by: Nguyen, Truc Mai-Thanh, et al.
Published: (2025)
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
by: Vo, Khang H. N., et al.
Published: (2025)
by: Vo, Khang H. N., et al.
Published: (2025)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
by: Van-Dinh, Tue-Thu, et al.
Published: (2025)
by: Van-Dinh, Tue-Thu, et al.
Published: (2025)
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking
by: Tran, Dien X., et al.
Published: (2025)
by: Tran, Dien X., et al.
Published: (2025)
RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
by: Nguyen, Long, et al.
Published: (2025)
by: Nguyen, Long, et al.
Published: (2025)
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
by: Nguyen, Long S. T., et al.
Published: (2025)
by: Nguyen, Long S. T., et al.
Published: (2025)
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
by: Nguyen, Ngoc-Son, et al.
Published: (2025)
by: Nguyen, Ngoc-Son, et al.
Published: (2025)
VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding
by: Do, Phong Nguyen-Thuan, et al.
Published: (2024)
by: Do, Phong Nguyen-Thuan, et al.
Published: (2024)
ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models
by: Nguyen, Duy Vu Minh, et al.
Published: (2026)
by: Nguyen, Duy Vu Minh, et al.
Published: (2026)
Describe Anything Model for Visual Question Answering on Text-rich Images
by: Vu, Yen-Linh, et al.
Published: (2025)
by: Vu, Yen-Linh, et al.
Published: (2025)
Similar Items
-
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
by: Pham, Huy Quang, et al.
Published: (2024) -
ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition
by: Nguyen, Khoa Anh, et al.
Published: (2026) -
ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images
by: Van Nguyen, Quan, et al.
Published: (2024) -
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026) -
Coreference Resolution for Vietnamese Narrative Texts
by: Tran, Hieu-Dai, et al.
Published: (2025)