KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
Fuente:
arXiv
Salvato in:
| Autori principali: | Pham, Anh-Cuong, Nguyen, Van-Quang, Vuong, Thi-Hong, Ha, Quang-Thuy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
di: Nguyen, Van Quang
Pubblicazione: (2026)
di: Nguyen, Van Quang
Pubblicazione: (2026)
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
di: Nguyen, Thi-Nhu-Quynh, et al.
Pubblicazione: (2024)
di: Nguyen, Thi-Nhu-Quynh, et al.
Pubblicazione: (2024)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
di: Nguyen, Quang-Binh, et al.
Pubblicazione: (2025)
di: Nguyen, Quang-Binh, et al.
Pubblicazione: (2025)
V-Math: An Agentic Approach to the Vietnamese National High School Graduation Mathematics Exams
di: Nguyen, Duong Q., et al.
Pubblicazione: (2025)
di: Nguyen, Duong Q., et al.
Pubblicazione: (2025)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
di: Nguyen, Toan, et al.
Pubblicazione: (2025)
di: Nguyen, Toan, et al.
Pubblicazione: (2025)
Stable Messenger: Steganography for Message-Concealed Image Generation
di: Nguyen, Quang, et al.
Pubblicazione: (2023)
di: Nguyen, Quang, et al.
Pubblicazione: (2023)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
di: Nguyen, Nhi Ngoc-Yen, et al.
Pubblicazione: (2026)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
di: Nguyen, Trong-Tung, et al.
Pubblicazione: (2024)
di: Nguyen, Trong-Tung, et al.
Pubblicazione: (2024)
UIT-DarkCow team at ImageCLEFmedical Caption 2024: Diagnostic Captioning for Radiology Images Efficiency with Transformer Models
di: Van Nguyen, Quan, et al.
Pubblicazione: (2024)
di: Van Nguyen, Quan, et al.
Pubblicazione: (2024)
LiteNeXt: A Novel Lightweight ConvMixer-based Model with Self-embedding Representation Parallel for Medical Image Segmentation
di: Tran, Ngoc-Du, et al.
Pubblicazione: (2024)
di: Tran, Ngoc-Du, et al.
Pubblicazione: (2024)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
di: Truong, Quang-Trung, et al.
Pubblicazione: (2025)
di: Truong, Quang-Trung, et al.
Pubblicazione: (2025)
SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities
di: Nguyen, Dung Thuy, et al.
Pubblicazione: (2025)
di: Nguyen, Dung Thuy, et al.
Pubblicazione: (2025)
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
di: Tran, Huyen T. T., et al.
Pubblicazione: (2026)
di: Tran, Huyen T. T., et al.
Pubblicazione: (2026)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
di: Tuong, Nguyen Anh, et al.
Pubblicazione: (2026)
di: Tuong, Nguyen Anh, et al.
Pubblicazione: (2026)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
di: Tran, Quoc-Khang, et al.
Pubblicazione: (2026)
di: Tran, Quoc-Khang, et al.
Pubblicazione: (2026)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
di: Dinh, Quang Minh, et al.
Pubblicazione: (2024)
di: Dinh, Quang Minh, et al.
Pubblicazione: (2024)
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
di: Nguyen, Phuc D. A., et al.
Pubblicazione: (2024)
di: Nguyen, Phuc D. A., et al.
Pubblicazione: (2024)
Virtual Fusion with Contrastive Learning for Single Sensor-based Activity Recognition
di: Nguyen, Duc-Anh, et al.
Pubblicazione: (2023)
di: Nguyen, Duc-Anh, et al.
Pubblicazione: (2023)
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
di: Nguyen, Hung, et al.
Pubblicazione: (2024)
di: Nguyen, Hung, et al.
Pubblicazione: (2024)
InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting
di: Vu, Duc, et al.
Pubblicazione: (2026)
di: Vu, Duc, et al.
Pubblicazione: (2026)
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
di: D., Quang-Anh N., et al.
Pubblicazione: (2024)
di: D., Quang-Anh N., et al.
Pubblicazione: (2024)
Comparing Deep Neural Network for Multi-Label ECG Diagnosis From Scanned ECG
di: Nguyen, Cuong V., et al.
Pubblicazione: (2025)
di: Nguyen, Cuong V., et al.
Pubblicazione: (2025)
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher
di: Dao, Trung, et al.
Pubblicazione: (2024)
di: Dao, Trung, et al.
Pubblicazione: (2024)
Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image Classification
di: Vu, Anh Mai, et al.
Pubblicazione: (2025)
di: Vu, Anh Mai, et al.
Pubblicazione: (2025)
AC-MAMBASEG: An adaptive convolution and Mamba-based architecture for enhanced skin lesion segmentation
di: Nguyen, Viet-Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Viet-Thanh, et al.
Pubblicazione: (2024)
SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation
di: Thai, Gia Huy, et al.
Pubblicazione: (2025)
di: Thai, Gia Huy, et al.
Pubblicazione: (2025)
FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
di: Pham, Trong Thang, et al.
Pubblicazione: (2024)
di: Pham, Trong Thang, et al.
Pubblicazione: (2024)
SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
di: Truong, Khang, et al.
Pubblicazione: (2025)
di: Truong, Khang, et al.
Pubblicazione: (2025)
PixLore: A Dataset-driven Approach to Rich Image Captioning
di: Bonilla-Salvador, Diego, et al.
Pubblicazione: (2023)
di: Bonilla-Salvador, Diego, et al.
Pubblicazione: (2023)
DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction
di: Van, Cuong Tran, et al.
Pubblicazione: (2026)
di: Van, Cuong Tran, et al.
Pubblicazione: (2026)
Region in Context: Text-condition Image editing with Human-like semantic reasoning
di: Vu, Thuy Phuong, et al.
Pubblicazione: (2025)
di: Vu, Thuy Phuong, et al.
Pubblicazione: (2025)
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
di: Huang, Yifeng, et al.
Pubblicazione: (2023)
di: Huang, Yifeng, et al.
Pubblicazione: (2023)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
di: Quang, Ngoc Bui Lam, et al.
Pubblicazione: (2025)
di: Quang, Ngoc Bui Lam, et al.
Pubblicazione: (2025)
MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment
di: Banerjee, Sagarika, et al.
Pubblicazione: (2026)
di: Banerjee, Sagarika, et al.
Pubblicazione: (2026)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
di: Vu, Huu-An, et al.
Pubblicazione: (2025)
di: Vu, Huu-An, et al.
Pubblicazione: (2025)
Image Captioning in news report scenario
di: Liu, Tianrui, et al.
Pubblicazione: (2024)
di: Liu, Tianrui, et al.
Pubblicazione: (2024)
SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
di: Nguyen, Kien, et al.
Pubblicazione: (2025)
di: Nguyen, Kien, et al.
Pubblicazione: (2025)
CaptionFool: Universal Image Captioning Model Attacks
di: Parekh, Swapnil
Pubblicazione: (2026)
di: Parekh, Swapnil
Pubblicazione: (2026)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
di: Pham, Trong-Thang, et al.
Pubblicazione: (2026)
di: Pham, Trong-Thang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
di: Pham, Huy Quang, et al.
Pubblicazione: (2024) -
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
di: Nguyen, Van Quang
Pubblicazione: (2026) -
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
di: Nguyen, Thi-Nhu-Quynh, et al.
Pubblicazione: (2024) -
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
di: Nguyen, Quang-Binh, et al.
Pubblicazione: (2025) -
V-Math: An Agentic Approach to the Vietnamese National High School Graduation Mathematics Exams
di: Nguyen, Duong Q., et al.
Pubblicazione: (2025)