BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Le, Huy, Chung, Nhat, Kieu, Tung, Nguyen, Anh, Le, Ngan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
von: Le, Huy, et al.
Veröffentlicht: (2023)
von: Le, Huy, et al.
Veröffentlicht: (2023)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
von: Le, Huy, et al.
Veröffentlicht: (2025)
von: Le, Huy, et al.
Veröffentlicht: (2025)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
von: Chung, Nhat, et al.
Veröffentlicht: (2025)
von: Chung, Nhat, et al.
Veröffentlicht: (2025)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
von: Nguyen, Long, et al.
Veröffentlicht: (2025)
von: Nguyen, Long, et al.
Veröffentlicht: (2025)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2024)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
von: Le, Tri, et al.
Veröffentlicht: (2025)
von: Le, Tri, et al.
Veröffentlicht: (2025)
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
von: Nguyen, Toan, et al.
Veröffentlicht: (2024)
von: Nguyen, Toan, et al.
Veröffentlicht: (2024)
Multimodal Contextualized Support for Enhancing Video Retrieval System
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
von: Nguyen, Ngoc Son, et al.
Veröffentlicht: (2024)
von: Nguyen, Ngoc Son, et al.
Veröffentlicht: (2024)
PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-Identification
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2023)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2023)
VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2026)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2026)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
von: Le, Minh, et al.
Veröffentlicht: (2025)
von: Le, Minh, et al.
Veröffentlicht: (2025)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
Learning Human Motion with Temporally Conditional Mamba
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
Link prediction Graph Neural Networks for structure recognition of Handwritten Mathematical Expressions
von: Nguyen, Cuong Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Cuong Tuan, et al.
Veröffentlicht: (2025)
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking
von: Tran, Huu-Loc, et al.
Veröffentlicht: (2025)
von: Tran, Huu-Loc, et al.
Veröffentlicht: (2025)
Towards Comprehensive Vietnamese Retrieval-Augmented Generation and Large Language Models
von: Duc, Nguyen Quang, et al.
Veröffentlicht: (2024)
von: Duc, Nguyen Quang, et al.
Veröffentlicht: (2024)
CLEAR: Causal Learning Framework For Robust Histopathology Tumor Detection Under Out-Of-Distribution Shifts
von: Thi, Kieu-Anh Truong, et al.
Veröffentlicht: (2025)
von: Thi, Kieu-Anh Truong, et al.
Veröffentlicht: (2025)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
von: Li, Qian, et al.
Veröffentlicht: (2024)
von: Li, Qian, et al.
Veröffentlicht: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Universal Multi-Domain Translation via Diffusion Routers
von: Kieu, Duc, et al.
Veröffentlicht: (2025)
von: Kieu, Duc, et al.
Veröffentlicht: (2025)
Studying and Mitigating Biases in Sign Language Understanding Models
von: Atwell, Katherine, et al.
Veröffentlicht: (2024)
von: Atwell, Katherine, et al.
Veröffentlicht: (2024)
HAtt-Flow: Hierarchical Attention-Flow Mechanism for Group Activity Scene Graph Generation in Videos
von: Chappa, Naga VS Raviteja, et al.
Veröffentlicht: (2023)
von: Chappa, Naga VS Raviteja, et al.
Veröffentlicht: (2023)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
von: Pham, Trong-Thang, et al.
Veröffentlicht: (2026)
von: Pham, Trong-Thang, et al.
Veröffentlicht: (2026)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
A Lightweight Moment Retrieval System with Global Re-Ranking and Robust Adaptive Bidirectional Temporal Search
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhu, Tinh-Anh, et al.
Veröffentlicht: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
von: Le, Nhat, et al.
Veröffentlicht: (2026)
von: Le, Nhat, et al.
Veröffentlicht: (2026)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
von: Vo, Khoa, et al.
Veröffentlicht: (2024)
von: Vo, Khoa, et al.
Veröffentlicht: (2024)
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2023)
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2023)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
von: Le, Huy, et al.
Veröffentlicht: (2023) -
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
von: Le, Huy, et al.
Veröffentlicht: (2025) -
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025) -
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
von: Chung, Nhat, et al.
Veröffentlicht: (2025) -
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)