A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions
Fuente:
arXiv
Salvato in:
| Autori principali: | Le, Anh, Lam, Thanh, Nguyen, Dung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
di: Nguyen, Huu Tien, et al.
Pubblicazione: (2025)
di: Nguyen, Huu Tien, et al.
Pubblicazione: (2025)
VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering
di: Nguyen, Hai-Dang, et al.
Pubblicazione: (2025)
di: Nguyen, Hai-Dang, et al.
Pubblicazione: (2025)
Improving Micro-Expression Recognition with Phase-Aware Temporal Augmentation
di: Khuong, Vu Tram Anh, et al.
Pubblicazione: (2025)
di: Khuong, Vu Tram Anh, et al.
Pubblicazione: (2025)
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
di: Nguyen, Khoi Anh, et al.
Pubblicazione: (2025)
di: Nguyen, Khoi Anh, et al.
Pubblicazione: (2025)
Apex-Centered Spatio-Temporal Rank Pooling and Gradient Attention for Micro-Expression Recognition
di: Nguyen, Luu Tu, et al.
Pubblicazione: (2025)
di: Nguyen, Luu Tu, et al.
Pubblicazione: (2025)
LaVy: Vietnamese Multimodal Large Language Model
di: Tran, Chi, et al.
Pubblicazione: (2024)
di: Tran, Chi, et al.
Pubblicazione: (2024)
Multi-view Action Recognition via Directed Gromov-Wasserstein Discrepancy
di: Nguyen, Hoang-Quan, et al.
Pubblicazione: (2024)
di: Nguyen, Hoang-Quan, et al.
Pubblicazione: (2024)
U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025
di: Le, Duc-Nhuan, et al.
Pubblicazione: (2026)
di: Le, Duc-Nhuan, et al.
Pubblicazione: (2026)
A Novel Combined Optical Flow Approach for Comprehensive Micro-Expression Recognition
di: Khuong, Vu Tram Anh, et al.
Pubblicazione: (2025)
di: Khuong, Vu Tram Anh, et al.
Pubblicazione: (2025)
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
di: Nguyen, Ngoc Son, et al.
Pubblicazione: (2024)
di: Nguyen, Ngoc Son, et al.
Pubblicazione: (2024)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
di: Tien, Dong Nguyen, et al.
Pubblicazione: (2025)
di: Tien, Dong Nguyen, et al.
Pubblicazione: (2025)
A Hybrid Vision Transformer Approach for Mathematical Expression Recognition
di: Le, Anh Duy, et al.
Pubblicazione: (2026)
di: Le, Anh Duy, et al.
Pubblicazione: (2026)
DIANet: A Phase-Aware Dual-Stream Network for Micro-Expression Recognition via Dynamic Images
di: Khuong, Vu Tram Anh, et al.
Pubblicazione: (2025)
di: Khuong, Vu Tram Anh, et al.
Pubblicazione: (2025)
Elderly Activity Recognition in the Wild: Results from the EAR Challenge
di: Duong, Anh-Kiet
Pubblicazione: (2025)
di: Duong, Anh-Kiet
Pubblicazione: (2025)
Place Recognition Meet Multiple Modalitie: A Comprehensive Review, Current Challenges and Future Directions
di: Li, Zhenyu, et al.
Pubblicazione: (2025)
di: Li, Zhenyu, et al.
Pubblicazione: (2025)
Dual-View Optical Flow for 4D Micro-Expression Recognition - A Multi-Stream Fusion Attention Approach
di: Nguyen, Luu Tu, et al.
Pubblicazione: (2026)
di: Nguyen, Luu Tu, et al.
Pubblicazione: (2026)
FMANet: A Novel Dual-Phase Optical Flow Approach with Fusion Motion Attention Network for Robust Micro-expression Recognition
di: Nguyen, Luu Tu, et al.
Pubblicazione: (2025)
di: Nguyen, Luu Tu, et al.
Pubblicazione: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
di: Tuong, Nguyen Anh, et al.
Pubblicazione: (2026)
di: Tuong, Nguyen Anh, et al.
Pubblicazione: (2026)
Adaptive Fusion Network with Temporal-Ranked and Motion-Intensity Dynamic Images for Micro-expression Recognition
di: Man, Thi Bich Phuong, et al.
Pubblicazione: (2025)
di: Man, Thi Bich Phuong, et al.
Pubblicazione: (2025)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2025)
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2025)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
di: Pham, Anh-Cuong, et al.
Pubblicazione: (2024)
di: Pham, Anh-Cuong, et al.
Pubblicazione: (2024)
Skeletal Video Anomaly Detection using Deep Learning: Survey, Challenges and Future Directions
di: Mishra, Pratik K., et al.
Pubblicazione: (2022)
di: Mishra, Pratik K., et al.
Pubblicazione: (2022)
Maximising the Utility of Validation Sets for Imbalanced Noisy-label Meta-learning
di: Hoang, Dung Anh, et al.
Pubblicazione: (2022)
di: Hoang, Dung Anh, et al.
Pubblicazione: (2022)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
di: Nguyen, Minh Khoi, et al.
Pubblicazione: (2026)
di: Nguyen, Minh Khoi, et al.
Pubblicazione: (2026)
Sharpness-Aware Data Generation for Zero-shot Quantization
di: Hoang-Anh, Dung, et al.
Pubblicazione: (2025)
di: Hoang-Anh, Dung, et al.
Pubblicazione: (2025)
Semantic Mapping in Indoor Embodied AI -- A Survey on Advances, Challenges, and Future Directions
di: Raychaudhuri, Sonia, et al.
Pubblicazione: (2025)
di: Raychaudhuri, Sonia, et al.
Pubblicazione: (2025)
MissBench: Benchmarking Multimodal Affective Analysis under Imbalanced Missing Modalities
di: Pham, Tien Anh, et al.
Pubblicazione: (2026)
di: Pham, Tien Anh, et al.
Pubblicazione: (2026)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
di: Nguyen, Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Hieu, et al.
Pubblicazione: (2024)
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
di: Le, Kha Nhat, et al.
Pubblicazione: (2024)
di: Le, Kha Nhat, et al.
Pubblicazione: (2024)
ACM Multimedia Grand Challenge on ENT Endoscopy Analysis
di: Nguyen, Trong-Thuan, et al.
Pubblicazione: (2025)
di: Nguyen, Trong-Thuan, et al.
Pubblicazione: (2025)
Visual Hand Gesture Recognition with Deep Learning: A Comprehensive Review of Methods, Datasets, Challenges and Future Research Directions
di: Foteinos, Konstantinos, et al.
Pubblicazione: (2025)
di: Foteinos, Konstantinos, et al.
Pubblicazione: (2025)
N-EIoU-YOLOv9: A Signal-Aware Bounding Box Regression Loss for Lightweight Mobile Detection of Rice Leaf Diseases
di: Duc, Dung Ta Nguyen, et al.
Pubblicazione: (2026)
di: Duc, Dung Ta Nguyen, et al.
Pubblicazione: (2026)
Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models
di: Hoang, Dung Anh, et al.
Pubblicazione: (2026)
di: Hoang, Dung Anh, et al.
Pubblicazione: (2026)
MetaAug: Meta-Data Augmentation for Post-Training Quantization
di: Pham, Cuong, et al.
Pubblicazione: (2024)
di: Pham, Cuong, et al.
Pubblicazione: (2024)
Driver Attention Tracking and Analysis
di: Nguyen, Dat Viet Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Dat Viet Thanh, et al.
Pubblicazione: (2024)
Emotic Masked Autoencoder with Attention Fusion for Facial Expression Recognition
di: Nguyen-Xuan, Bach, et al.
Pubblicazione: (2024)
di: Nguyen-Xuan, Bach, et al.
Pubblicazione: (2024)
DiffAugment: Diffusion based Long-Tailed Visual Relationship Recognition
di: Gupta, Parul, et al.
Pubblicazione: (2024)
di: Gupta, Parul, et al.
Pubblicazione: (2024)
PADM: A Physics-aware Diffusion Model for Attenuation Correction
di: Pham, Trung Kien, et al.
Pubblicazione: (2025)
di: Pham, Trung Kien, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
di: Nguyen, Huu Tien, et al.
Pubblicazione: (2025) -
VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering
di: Nguyen, Hai-Dang, et al.
Pubblicazione: (2025) -
Improving Micro-Expression Recognition with Phase-Aware Temporal Augmentation
di: Khuong, Vu Tram Anh, et al.
Pubblicazione: (2025) -
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
di: Nguyen, Khoi Anh, et al.
Pubblicazione: (2025) -
Apex-Centered Spatio-Temporal Rank Pooling and Gradient Attention for Micro-Expression Recognition
di: Nguyen, Luu Tu, et al.
Pubblicazione: (2025)