Multimodal Contextualized Support for Enhancing Video Retrieval System
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen-Le, Quoc-Bao, Le-Nguyen, Thanh-Huy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
por: Vu, Huu-An, et al.
Publicado: (2025)
por: Vu, Huu-An, et al.
Publicado: (2025)
RotCAtt-TransUNet++: Novel Deep Neural Network for Sophisticated Cardiac Segmentation
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
por: Le, Huy, et al.
Publicado: (2023)
por: Le, Huy, et al.
Publicado: (2023)
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
por: Luu, Vinh Quoc, et al.
Publicado: (2024)
por: Luu, Vinh Quoc, et al.
Publicado: (2024)
Examining Monitoring System: Detecting Abnormal Behavior In Online Examinations
por: Ngo, Dinh An, et al.
Publicado: (2024)
por: Ngo, Dinh An, et al.
Publicado: (2024)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
por: Le, Huy, et al.
Publicado: (2025)
por: Le, Huy, et al.
Publicado: (2025)
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
por: Nguyen, Cong-Duy, et al.
Publicado: (2025)
por: Nguyen, Cong-Duy, et al.
Publicado: (2025)
Enhancing the Fairness and Performance of Edge Cameras with Explainable AI
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
por: Quang, Ngoc Bui Lam, et al.
Publicado: (2025)
por: Quang, Ngoc Bui Lam, et al.
Publicado: (2025)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
por: Phung, Thu Hang, et al.
Publicado: (2026)
por: Phung, Thu Hang, et al.
Publicado: (2026)
Novel 3D Binary Indexed Tree for Volume Computation of 3D Reconstructed Models from Volumetric Data
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
por: Thanh, Toan Le Ngo, et al.
Publicado: (2025)
por: Thanh, Toan Le Ngo, et al.
Publicado: (2025)
Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework
por: Nguyen, Cong Huy, et al.
Publicado: (2026)
por: Nguyen, Cong Huy, et al.
Publicado: (2026)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
por: Nguyen-Nhu, Tinh-Anh, et al.
Publicado: (2025)
por: Nguyen-Nhu, Tinh-Anh, et al.
Publicado: (2025)
Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image Classification
por: Vu, Anh Mai, et al.
Publicado: (2025)
por: Vu, Anh Mai, et al.
Publicado: (2025)
HDC: Hierarchical Distillation for Multi-level Noisy Consistency in Semi-Supervised Fetal Ultrasound Segmentation
por: Le, Tran Quoc Khanh, et al.
Publicado: (2025)
por: Le, Tran Quoc Khanh, et al.
Publicado: (2025)
IGL-DT: Iterative Global-Local Feature Learning with Dual-Teacher Semantic Segmentation Framework under Limited Annotation Scheme
por: Tran, Dinh Dai Quan, et al.
Publicado: (2025)
por: Tran, Dinh Dai Quan, et al.
Publicado: (2025)
Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models
por: Le, Huy Hoan, et al.
Publicado: (2025)
por: Le, Huy Hoan, et al.
Publicado: (2025)
Lightweight Models for Emotional Analysis in Video
por: Nguyen, Quoc-Tien, et al.
Publicado: (2025)
por: Nguyen, Quoc-Tien, et al.
Publicado: (2025)
Learning Disentangled Stain and Structural Representations for Semi-Supervised Histopathology Segmentation
por: Pham, Ha-Hieu, et al.
Publicado: (2025)
por: Pham, Ha-Hieu, et al.
Publicado: (2025)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
por: Nguyen, Toan, et al.
Publicado: (2025)
por: Nguyen, Toan, et al.
Publicado: (2025)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
por: Nguyen, Tien-Huy, et al.
Publicado: (2026)
por: Nguyen, Tien-Huy, et al.
Publicado: (2026)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
por: Tran, Quoc-Khang, et al.
Publicado: (2026)
por: Tran, Quoc-Khang, et al.
Publicado: (2026)
Brain Tumor Segmentation in MRI Images with 3D U-Net and Contextual Transformer
por: Nguyen, Thien-Qua T., et al.
Publicado: (2024)
por: Nguyen, Thien-Qua T., et al.
Publicado: (2024)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
por: Tuong, Nguyen Anh, et al.
Publicado: (2026)
por: Tuong, Nguyen Anh, et al.
Publicado: (2026)
Efficient and Concise Explanations for Object Detection with Gaussian-Class Activation Mapping Explainer
por: Nguyen, Quoc Khanh, et al.
Publicado: (2024)
por: Nguyen, Quoc Khanh, et al.
Publicado: (2024)
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
por: Pham, Trong Thang, et al.
Publicado: (2026)
por: Pham, Trong Thang, et al.
Publicado: (2026)
MasHeNe: A Benchmark for Head and Neck CT Mass Segmentation using Window-Enhanced Mamba with Frequency-Domain Integration
por: Dao, Thao Thi Phuong, et al.
Publicado: (2025)
por: Dao, Thao Thi Phuong, et al.
Publicado: (2025)
Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics
por: Le, Minh H. N., et al.
Publicado: (2026)
por: Le, Minh H. N., et al.
Publicado: (2026)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
por: Le, Huy, et al.
Publicado: (2025)
por: Le, Huy, et al.
Publicado: (2025)
LightX3ECG: A Lightweight and eXplainable Deep Learning System for 3-lead Electrocardiogram Classification
por: Le, Khiem H., et al.
Publicado: (2022)
por: Le, Khiem H., et al.
Publicado: (2022)
ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising
por: Chaubey, Ashutosh, et al.
Publicado: (2024)
por: Chaubey, Ashutosh, et al.
Publicado: (2024)
Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
por: Shalabi, Fatma, et al.
Publicado: (2024)
por: Shalabi, Fatma, et al.
Publicado: (2024)
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher
por: Dao, Trung, et al.
Publicado: (2024)
por: Dao, Trung, et al.
Publicado: (2024)
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
por: Nguyen, Thanh Tam, et al.
Publicado: (2024)
por: Nguyen, Thanh Tam, et al.
Publicado: (2024)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
por: Pham, Trong-Thang, et al.
Publicado: (2026)
por: Pham, Trong-Thang, et al.
Publicado: (2026)
From Specialist to Generalist: Unlocking SAM's Learning Potential on Unlabeled Medical Images
por: Vu, Vi, et al.
Publicado: (2026)
por: Vu, Vi, et al.
Publicado: (2026)
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
por: Nguyen, Hung, et al.
Publicado: (2024)
por: Nguyen, Hung, et al.
Publicado: (2024)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
por: Nguyen, Quang Vinh, et al.
Publicado: (2024)
por: Nguyen, Quang Vinh, et al.
Publicado: (2024)
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
Ejemplares similares
-
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
por: Vu, Huu-An, et al.
Publicado: (2025) -
RotCAtt-TransUNet++: Novel Deep Neural Network for Sophisticated Cardiac Segmentation
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024) -
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
por: Le, Huy, et al.
Publicado: (2023) -
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
por: Luu, Vinh Quoc, et al.
Publicado: (2024) -
Examining Monitoring System: Detecting Abnormal Behavior In Online Examinations
por: Ngo, Dinh An, et al.
Publicado: (2024)