SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Ruixiang, Xu, Zhihao, Lan, Bangxiang, Xin, Zijie, Liu, Jingyu, Li, Xirong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Object Sketch Animation by Scene Decomposition and Motion Planning
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
by: Lan, Bangxiang, et al.
Published: (2025)
by: Lan, Bangxiang, et al.
Published: (2025)
Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data
by: Deng, Yuchuan, et al.
Published: (2026)
by: Deng, Yuchuan, et al.
Published: (2026)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
by: Hu, Fan, et al.
Published: (2025)
by: Hu, Fan, et al.
Published: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
by: Zhao, Ruixiang, et al.
Published: (2024)
by: Zhao, Ruixiang, et al.
Published: (2024)
Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
EA-VTR: Event-Aware Video-Text Retrieval
by: Ma, Zongyang, et al.
Published: (2024)
by: Ma, Zongyang, et al.
Published: (2024)
Adversarial Video Promotion Against Text-to-Video Retrieval
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
PRVR: Partially Relevant Video Retrieval
by: Chen, Xianke, et al.
Published: (2022)
by: Chen, Xianke, et al.
Published: (2022)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
by: Dou, Huanzhang, et al.
Published: (2024)
by: Dou, Huanzhang, et al.
Published: (2024)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
by: Tian, Kaibin, et al.
Published: (2024)
by: Tian, Kaibin, et al.
Published: (2024)
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
by: Fu, Yuhan, et al.
Published: (2024)
by: Fu, Yuhan, et al.
Published: (2024)
HVD: Human Vision-Driven Video Representation Learning for Text-Video Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
by: Zhao, Zecheng, et al.
Published: (2025)
by: Zhao, Zecheng, et al.
Published: (2025)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
by: Cho, CH, et al.
Published: (2025)
by: Cho, CH, et al.
Published: (2025)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
by: Liu, Weijia, et al.
Published: (2025)
by: Liu, Weijia, et al.
Published: (2025)
Co-Teaching for Unsupervised Domain Adaptation and Expansion
by: Lin, Hailan, et al.
Published: (2022)
by: Lin, Hailan, et al.
Published: (2022)
BADet: Boundary-Aware 3D Object Detection from Point Clouds
by: Qian, Rui, et al.
Published: (2021)
by: Qian, Rui, et al.
Published: (2021)
D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching
by: Liu, Jingyu, et al.
Published: (2024)
by: Liu, Jingyu, et al.
Published: (2024)
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
by: Shen, Leqi, et al.
Published: (2024)
by: Shen, Leqi, et al.
Published: (2024)
T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
by: Li, Yili, et al.
Published: (2024)
by: Li, Yili, et al.
Published: (2024)
TokenBinder: Text-Video Retrieval with One-to-Many Alignment Paradigm
by: Zhang, Bingqing, et al.
Published: (2024)
by: Zhang, Bingqing, et al.
Published: (2024)
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
by: Li, Wenjun, et al.
Published: (2024)
by: Li, Wenjun, et al.
Published: (2024)
Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models
by: Cao, Cong, et al.
Published: (2026)
by: Cao, Cong, et al.
Published: (2026)
EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval
by: Chen, Yuhan, et al.
Published: (2026)
by: Chen, Yuhan, et al.
Published: (2026)
Reinforcing Consistency in Video MLLMs with Structured Rewards
by: Quan, Yihao, et al.
Published: (2026)
by: Quan, Yihao, et al.
Published: (2026)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2025)
by: Xiao, Jian, et al.
Published: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
TVPR: Text-to-Video Person Retrieval and a New Benchmark
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
by: Wang, Jiamian, et al.
Published: (2024)
by: Wang, Jiamian, et al.
Published: (2024)
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
by: Lou, Zijie, et al.
Published: (2026)
by: Lou, Zijie, et al.
Published: (2026)
View while Moving: Efficient Video Recognition in Long-untrimmed Videos
by: Tian, Ye, et al.
Published: (2023)
by: Tian, Ye, et al.
Published: (2023)
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
by: Shi, Miaojing, et al.
Published: (2026)
by: Shi, Miaojing, et al.
Published: (2026)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024)
by: Zhang, Haonan, et al.
Published: (2024)
Similar Items
-
Multi-Object Sketch Animation by Scene Decomposition and Motion Planning
by: Liu, Jingyu, et al.
Published: (2025) -
Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
by: Lan, Bangxiang, et al.
Published: (2025) -
Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data
by: Deng, Yuchuan, et al.
Published: (2026) -
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026) -
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
by: Hu, Fan, et al.
Published: (2025)