Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Hieu, Ta, Cong-Hoang, Le-Nguyen, Phuong-Thuy, Tran, Minh-Triet, Le, Trung-Nghia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
TaleForge: Interactive Multimodal System for Personalized Story Creation
di: Nguyen, Minh-Loi, et al.
Pubblicazione: (2025)
di: Nguyen, Minh-Loi, et al.
Pubblicazione: (2025)
GUNNEL: Guided Mixup Augmentation and Multi-Model Fusion for Aquatic Animal Segmentation
di: Le, Minh-Quan, et al.
Pubblicazione: (2021)
di: Le, Minh-Quan, et al.
Pubblicazione: (2021)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
GenFlow: Interactive Modular System for Image Generation
di: Nguyen, Duc-Hung, et al.
Pubblicazione: (2025)
di: Nguyen, Duc-Hung, et al.
Pubblicazione: (2025)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
CamoFA: A Learnable Fourier-based Augmentation for Camouflage Segmentation
di: Le, Minh-Quan, et al.
Pubblicazione: (2023)
di: Le, Minh-Quan, et al.
Pubblicazione: (2023)
GenKOL: Modular Generative AI Framework For Scalable Virtual KOL Generation
di: To, Tan-Hiep, et al.
Pubblicazione: (2025)
di: To, Tan-Hiep, et al.
Pubblicazione: (2025)
Automated Image Recognition Framework
di: Nguyen, Quang-Binh, et al.
Pubblicazione: (2025)
di: Nguyen, Quang-Binh, et al.
Pubblicazione: (2025)
Interactive Interface For Semantic Segmentation Dataset Synthesis
di: Tran, Ngoc-Do, et al.
Pubblicazione: (2025)
di: Tran, Ngoc-Do, et al.
Pubblicazione: (2025)
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2026)
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2026)
ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation
di: Hoang, Trong-Vu, et al.
Pubblicazione: (2025)
di: Hoang, Trong-Vu, et al.
Pubblicazione: (2025)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation
di: Le, Minh-Quan, et al.
Pubblicazione: (2023)
di: Le, Minh-Quan, et al.
Pubblicazione: (2023)
ARtVista: Gateway To Empower Anyone Into Artist
di: Hoang, Trong-Vu, et al.
Pubblicazione: (2024)
di: Hoang, Trong-Vu, et al.
Pubblicazione: (2024)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
KiseKloset for Fashion Retrieval and Recommendation
di: Phan-Nguyen, Thanh-Tung, et al.
Pubblicazione: (2025)
di: Phan-Nguyen, Thanh-Tung, et al.
Pubblicazione: (2025)
Retrospective Feature Estimation for Continual Learning
di: Nguyen, Nghia D., et al.
Pubblicazione: (2024)
di: Nguyen, Nghia D., et al.
Pubblicazione: (2024)
FaR: Enhancing Multi-Concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
di: Tran, Gia-Nghia, et al.
Pubblicazione: (2025)
di: Tran, Gia-Nghia, et al.
Pubblicazione: (2025)
VisionGuard: Synergistic Framework for Helmet Violation Detection
di: Nguyen, Lam-Huy, et al.
Pubblicazione: (2025)
di: Nguyen, Lam-Huy, et al.
Pubblicazione: (2025)
ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
di: Nguyen, Thinh-Phuc, et al.
Pubblicazione: (2025)
di: Nguyen, Thinh-Phuc, et al.
Pubblicazione: (2025)
MasHeNe: A Benchmark for Head and Neck CT Mass Segmentation using Window-Enhanced Mamba with Frequency-Domain Integration
di: Dao, Thao Thi Phuong, et al.
Pubblicazione: (2025)
di: Dao, Thao Thi Phuong, et al.
Pubblicazione: (2025)
CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
Efficient 3D Brain Tumor Segmentation with Axial-Coronal-Sagittal Embedding
di: Huynh, Tuan-Luc, et al.
Pubblicazione: (2025)
di: Huynh, Tuan-Luc, et al.
Pubblicazione: (2025)
Enhancing Video Summarization with Context Awareness
di: Huynh-Lam, Hai-Dang, et al.
Pubblicazione: (2024)
di: Huynh-Lam, Hai-Dang, et al.
Pubblicazione: (2024)
Cluster-based Video Summarization with Temporal Context Awareness
di: Huynh-Lam, Hai-Dang, et al.
Pubblicazione: (2024)
di: Huynh-Lam, Hai-Dang, et al.
Pubblicazione: (2024)
PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback
di: Bui, Duy-Bao, et al.
Pubblicazione: (2025)
di: Bui, Duy-Bao, et al.
Pubblicazione: (2025)
iCONTRA: Toward Thematic Collection Design Via Interactive Concept Transfer
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2024)
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2024)
Shape2Animal: Creative Animal Generation from Natural Silhouettes
di: Tran, Quoc-Duy, et al.
Pubblicazione: (2025)
di: Tran, Quoc-Duy, et al.
Pubblicazione: (2025)
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
CAKE: Real-time Action Detection via Motion Distillation and Background-aware Contrastive Learning
di: Hoang, Hieu, et al.
Pubblicazione: (2026)
di: Hoang, Hieu, et al.
Pubblicazione: (2026)
Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025
di: Tran, Thien-Phuc, et al.
Pubblicazione: (2025)
di: Tran, Thien-Phuc, et al.
Pubblicazione: (2025)
Toward Content-based Indexing and Retrieval of Head and Neck CT with Abscess Segmentation
di: Dao, Thao Thi Phuong, et al.
Pubblicazione: (2025)
di: Dao, Thao Thi Phuong, et al.
Pubblicazione: (2025)
Learning to Stop Overthinking at Test Time
di: Bao, Hieu Tran, et al.
Pubblicazione: (2025)
di: Bao, Hieu Tran, et al.
Pubblicazione: (2025)
N-EIoU-YOLOv9: A Signal-Aware Bounding Box Regression Loss for Lightweight Mobile Detection of Rice Leaf Diseases
di: Duc, Dung Ta Nguyen, et al.
Pubblicazione: (2026)
di: Duc, Dung Ta Nguyen, et al.
Pubblicazione: (2026)
Evaluating Precise Geolocation Inference Capabilities of Vision Language Models
di: Jay, Neel, et al.
Pubblicazione: (2025)
di: Jay, Neel, et al.
Pubblicazione: (2025)
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
di: Nguyen, Huu Tien, et al.
Pubblicazione: (2025)
di: Nguyen, Huu Tien, et al.
Pubblicazione: (2025)
SATURN: Autoregressive Image Generation Guided by Scene Graphs
di: Vo, Thanh-Nhan, et al.
Pubblicazione: (2025)
di: Vo, Thanh-Nhan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024) -
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
di: Nguyen, Hieu, et al.
Pubblicazione: (2025) -
TaleForge: Interactive Multimodal System for Personalized Story Creation
di: Nguyen, Minh-Loi, et al.
Pubblicazione: (2025) -
GUNNEL: Guided Mixup Augmentation and Multi-Model Fusion for Aquatic Animal Segmentation
di: Le, Minh-Quan, et al.
Pubblicazione: (2021) -
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)