WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
Fuente:
arXiv
Salvato in:
| Autori principali: | Le, Huy, Kieu, Tung, Nguyen, Anh, Le, Ngan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
di: Le, Huy, et al.
Pubblicazione: (2025)
di: Le, Huy, et al.
Pubblicazione: (2025)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
di: Le, Huy, et al.
Pubblicazione: (2025)
di: Le, Huy, et al.
Pubblicazione: (2025)
Multimodal Contextualized Support for Enhancing Video Retrieval System
di: Nguyen-Le, Quoc-Bao, et al.
Pubblicazione: (2024)
di: Nguyen-Le, Quoc-Bao, et al.
Pubblicazione: (2024)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
di: Hanyu, Taisei, et al.
Pubblicazione: (2025)
di: Hanyu, Taisei, et al.
Pubblicazione: (2025)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
di: Chung, Nhat, et al.
Pubblicazione: (2025)
di: Chung, Nhat, et al.
Pubblicazione: (2025)
RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
di: Nguyen, Minh Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Minh Anh, et al.
Pubblicazione: (2026)
CLEAR: Causal Learning Framework For Robust Histopathology Tumor Detection Under Out-Of-Distribution Shifts
di: Thi, Kieu-Anh Truong, et al.
Pubblicazione: (2025)
di: Thi, Kieu-Anh Truong, et al.
Pubblicazione: (2025)
ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection
di: Nguyen, Minh Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Minh Anh, et al.
Pubblicazione: (2026)
Universal Multi-Domain Translation via Diffusion Routers
di: Kieu, Duc, et al.
Pubblicazione: (2025)
di: Kieu, Duc, et al.
Pubblicazione: (2025)
Adaptive Event Stream Slicing for Open-Vocabulary Event-Based Object Detection via Vision-Language Knowledge Distillation
di: Zhang, Jinchang, et al.
Pubblicazione: (2025)
di: Zhang, Jinchang, et al.
Pubblicazione: (2025)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
di: Pham, Trong-Thang, et al.
Pubblicazione: (2026)
di: Pham, Trong-Thang, et al.
Pubblicazione: (2026)
BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Rates
di: Nguyen, Phuong-Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Phuong-Anh, et al.
Pubblicazione: (2026)
Diversity-Aware Agnostic Ensemble of Sharpness Minimizers
di: Bui, Anh, et al.
Pubblicazione: (2024)
di: Bui, Anh, et al.
Pubblicazione: (2024)
Frequency Attention for Knowledge Distillation
di: Pham, Cuong, et al.
Pubblicazione: (2024)
di: Pham, Cuong, et al.
Pubblicazione: (2024)
VisionKG: Unleashing the Power of Visual Datasets via Knowledge Graph
di: Yuan, Jicheng, et al.
Pubblicazione: (2023)
di: Yuan, Jicheng, et al.
Pubblicazione: (2023)
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
WriteViT: Handwritten Text Generation with Vision Transformer
di: Nam, Dang Hoai, et al.
Pubblicazione: (2025)
di: Nam, Dang Hoai, et al.
Pubblicazione: (2025)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
di: Le, Tri, et al.
Pubblicazione: (2025)
di: Le, Tri, et al.
Pubblicazione: (2025)
FA-Seg: A Fast and Accurate Diffusion-Based Method for Open-Vocabulary Segmentation
di: Che, Huy, et al.
Pubblicazione: (2025)
di: Che, Huy, et al.
Pubblicazione: (2025)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
di: Quang, Ngoc Bui Lam, et al.
Pubblicazione: (2025)
di: Quang, Ngoc Bui Lam, et al.
Pubblicazione: (2025)
SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-world Object Detector
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding
di: Pham, Trong Thang, et al.
Pubblicazione: (2026)
di: Pham, Trong Thang, et al.
Pubblicazione: (2026)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
di: Nguyen-Nhu, Tinh-Anh, et al.
Pubblicazione: (2025)
di: Nguyen-Nhu, Tinh-Anh, et al.
Pubblicazione: (2025)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
di: Vo, Hao, et al.
Pubblicazione: (2026)
di: Vo, Hao, et al.
Pubblicazione: (2026)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
di: Vo, Khoa, et al.
Pubblicazione: (2024)
di: Vo, Khoa, et al.
Pubblicazione: (2024)
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
di: Rahmanzadehgervi, Pooyan, et al.
Pubblicazione: (2024)
di: Rahmanzadehgervi, Pooyan, et al.
Pubblicazione: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
di: Le, Quang-Hung, et al.
Pubblicazione: (2024)
di: Le, Quang-Hung, et al.
Pubblicazione: (2024)
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking
di: Tran, Huu-Loc, et al.
Pubblicazione: (2025)
di: Tran, Huu-Loc, et al.
Pubblicazione: (2025)
Learning Human Motion with Temporally Conditional Mamba
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
di: Nguyen, Minh Khoi, et al.
Pubblicazione: (2026)
di: Nguyen, Minh Khoi, et al.
Pubblicazione: (2026)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
TP-GMOT: Tracking Generic Multiple Object by Textual Prompt with Motion-Appearance Cost (MAC) SORT
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
di: Le, Minh, et al.
Pubblicazione: (2025)
di: Le, Minh, et al.
Pubblicazione: (2025)
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
di: Nguyen, Thuan Hoang, et al.
Pubblicazione: (2023)
di: Nguyen, Thuan Hoang, et al.
Pubblicazione: (2023)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
di: Wu, Size, et al.
Pubblicazione: (2023)
di: Wu, Size, et al.
Pubblicazione: (2023)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
di: Nguyen, Trong-Tung, et al.
Pubblicazione: (2024)
di: Nguyen, Trong-Tung, et al.
Pubblicazione: (2024)
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
di: Truong, Bao, et al.
Pubblicazione: (2026)
di: Truong, Bao, et al.
Pubblicazione: (2026)
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
Documenti analoghi
-
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
di: Le, Huy, et al.
Pubblicazione: (2025) -
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
di: Le, Huy, et al.
Pubblicazione: (2025) -
Multimodal Contextualized Support for Enhancing Video Retrieval System
di: Nguyen-Le, Quoc-Bao, et al.
Pubblicazione: (2024) -
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
di: Hanyu, Taisei, et al.
Pubblicazione: (2025) -
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
di: Chung, Nhat, et al.
Pubblicazione: (2025)