ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tran, Quoc-Khang, Nguyen, Minh-Thien, Pham, Nguyen-Khang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Model Soups to Classify Intangible Cultural Heritage Images from the Mekong Delta
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
A Vision-Language Foundation Model for Leaf Disease Identification
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2025)
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2025)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2026)
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2026)
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
von: Nguyen, Huu Tien, et al.
Veröffentlicht: (2025)
von: Nguyen, Huu Tien, et al.
Veröffentlicht: (2025)
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
von: Nguyen, Trong-Thuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Trong-Thuan, et al.
Veröffentlicht: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
von: Van-Dinh, Tue-Thu, et al.
Veröffentlicht: (2025)
von: Van-Dinh, Tue-Thu, et al.
Veröffentlicht: (2025)
PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Distortion-Aware Adversarial Attacks on Bounding Boxes of Object Detectors
von: Phuc, Pham, et al.
Veröffentlicht: (2024)
von: Phuc, Pham, et al.
Veröffentlicht: (2024)
Lifelong Whole Slide Image Analysis: Online Vision-Language Adaptation and Past-to-Present Gradient Distillation
von: Bui, Doanh C., et al.
Veröffentlicht: (2025)
von: Bui, Doanh C., et al.
Veröffentlicht: (2025)
GenKOL: Modular Generative AI Framework For Scalable Virtual KOL Generation
von: To, Tan-Hiep, et al.
Veröffentlicht: (2025)
von: To, Tan-Hiep, et al.
Veröffentlicht: (2025)
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
von: Vo, Khang H. N., et al.
Veröffentlicht: (2025)
von: Vo, Khang H. N., et al.
Veröffentlicht: (2025)
Enhancing YOLOv11n for Reliable Child Detection in Noisy Surveillance Footage
von: Tran, Khanh Linh, et al.
Veröffentlicht: (2026)
von: Tran, Khanh Linh, et al.
Veröffentlicht: (2026)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
von: Nguyen, Viet, et al.
Veröffentlicht: (2024)
von: Nguyen, Viet, et al.
Veröffentlicht: (2024)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
von: Nguyen, Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2024)
MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
von: Bui, Doanh C., et al.
Veröffentlicht: (2025)
von: Bui, Doanh C., et al.
Veröffentlicht: (2025)
A Lightweight Moment Retrieval System with Global Re-Ranking and Robust Adaptive Bidirectional Temporal Search
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
Volumetric Mapping with Panoptic Refinement via Kernel Density Estimation for Mobile Robots
von: Nguyen, Khang, et al.
Veröffentlicht: (2024)
von: Nguyen, Khang, et al.
Veröffentlicht: (2024)
CovHuSeg: An Enhanced Approach for Kidney Pathology Segmentation
von: Trinh, Huy, et al.
Veröffentlicht: (2024)
von: Trinh, Huy, et al.
Veröffentlicht: (2024)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
von: Le, Tri, et al.
Veröffentlicht: (2025)
von: Le, Tri, et al.
Veröffentlicht: (2025)
SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
von: Nguyen, Kien, et al.
Veröffentlicht: (2025)
von: Nguyen, Kien, et al.
Veröffentlicht: (2025)
QMaxViT-Unet+: A Query-Based MaxViT-Unet with Edge Enhancement for Scribble-Supervised Segmentation of Medical Images
von: Nguyen-Tat, Thien B., et al.
Veröffentlicht: (2025)
von: Nguyen-Tat, Thien B., et al.
Veröffentlicht: (2025)
V3D-SLAM: Robust RGB-D SLAM in Dynamic Environments with 3D Semantic Geometry Voting
von: Dang, Tuan, et al.
Veröffentlicht: (2024)
von: Dang, Tuan, et al.
Veröffentlicht: (2024)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
Enhancing the Fairness and Performance of Edge Cameras with Explainable AI
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
von: Pham, Duc-Hai, et al.
Veröffentlicht: (2024)
von: Pham, Duc-Hai, et al.
Veröffentlicht: (2024)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
von: Pham, Hieu Dinh Trung, et al.
Veröffentlicht: (2025)
von: Pham, Hieu Dinh Trung, et al.
Veröffentlicht: (2025)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
UIT-DarkCow team at ImageCLEFmedical Caption 2024: Diagnostic Captioning for Radiology Images Efficiency with Transformer Models
von: Van Nguyen, Quan, et al.
Veröffentlicht: (2024)
von: Van Nguyen, Quan, et al.
Veröffentlicht: (2024)
Neural Geometry Image-Based Representations with Optimal Transport (OT)
von: Gao, Xiang, et al.
Veröffentlicht: (2025)
von: Gao, Xiang, et al.
Veröffentlicht: (2025)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown Domains
von: Pham, Bang-Dang, et al.
Veröffentlicht: (2024)
von: Pham, Bang-Dang, et al.
Veröffentlicht: (2024)
PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback
von: Bui, Duy-Bao, et al.
Veröffentlicht: (2025)
von: Bui, Duy-Bao, et al.
Veröffentlicht: (2025)
SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
von: Truong, Khang, et al.
Veröffentlicht: (2025)
von: Truong, Khang, et al.
Veröffentlicht: (2025)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
von: Quang, Ngoc Bui Lam, et al.
Veröffentlicht: (2025)
von: Quang, Ngoc Bui Lam, et al.
Veröffentlicht: (2025)
LGCA: Enhancing Semantic Representation via Progressive Expansion
von: Cao, Thanh Hieu, et al.
Veröffentlicht: (2025)
von: Cao, Thanh Hieu, et al.
Veröffentlicht: (2025)
DualProtoSeg: Simple and Efficient Design with Text- and Image-Guided Prototype Learning for Weakly Supervised Histopathology Image Segmentation
von: Vu, Anh M., et al.
Veröffentlicht: (2025)
von: Vu, Anh M., et al.
Veröffentlicht: (2025)
MedSAD-CLIP: Supervised CLIP with Token-Patch Cross-Attention for Medical Anomaly Detection and Segmentation
von: Tran, Thuy Truong, et al.
Veröffentlicht: (2026)
von: Tran, Thuy Truong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Leveraging Model Soups to Classify Intangible Cultural Heritage Images from the Mekong Delta
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026) -
A Vision-Language Foundation Model for Leaf Disease Identification
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2025) -
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024) -
LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
von: Quoc, Khang Nguyen, et al.
Veröffentlicht: (2026) -
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
von: Nguyen, Huu Tien, et al.
Veröffentlicht: (2025)