Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zecheng, Chen, Zhi, Huang, Zi, Sadiq, Shazia, Chen, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
by: Zhao, Zecheng, et al.
Published: (2026)
by: Zhao, Zecheng, et al.
Published: (2026)
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos
by: Zhao, Zecheng, et al.
Published: (2025)
by: Zhao, Zecheng, et al.
Published: (2025)
FastEdit: Fast Text-Guided Single-Image Editing via Semantic-Aware Diffusion Fine-Tuning
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
by: Zhang, Deyu, et al.
Published: (2025)
by: Zhang, Deyu, et al.
Published: (2025)
EA-VTR: Event-Aware Video-Text Retrieval
by: Ma, Zongyang, et al.
Published: (2024)
by: Ma, Zongyang, et al.
Published: (2024)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
Motion-Aware Video Frame Interpolation
by: Han, Pengfei, et al.
Published: (2024)
by: Han, Pengfei, et al.
Published: (2024)
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
by: Hur, Chan, et al.
Published: (2025)
by: Hur, Chan, et al.
Published: (2025)
TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation
by: Lyu, Zonglin, et al.
Published: (2025)
by: Lyu, Zonglin, et al.
Published: (2025)
Shot-Aware Frame Sampling for Video Understanding
by: Zhao, Mengyu, et al.
Published: (2026)
by: Zhao, Mengyu, et al.
Published: (2026)
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
by: Liu, Yunze, et al.
Published: (2026)
by: Liu, Yunze, et al.
Published: (2026)
EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval
by: Chen, Yuhan, et al.
Published: (2026)
by: Chen, Yuhan, et al.
Published: (2026)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
Anatomy-Aware Conditional Image-Text Retrieval
by: Zheng, Meng, et al.
Published: (2025)
by: Zheng, Meng, et al.
Published: (2025)
Adversarial Video Promotion Against Text-to-Video Retrieval
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking
by: Zhao, Zixiang, et al.
Published: (2025)
by: Zhao, Zixiang, et al.
Published: (2025)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
CF-PRNet: Coarse-to-Fine Prototype Refining Network for Point Cloud Completion and Reconstruction
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
by: Wang, Shaokun, et al.
Published: (2026)
by: Wang, Shaokun, et al.
Published: (2026)
GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection
by: Chae, Joongwon, et al.
Published: (2026)
by: Chae, Joongwon, et al.
Published: (2026)
Bring Your Dreams to Life: Continual Text-to-Video Customization
by: Dong, Jiahua, et al.
Published: (2025)
by: Dong, Jiahua, et al.
Published: (2025)
AutoLoRA: Automatic LoRA Retrieval and Fine-Grained Gated Fusion for Text-to-Image Generation
by: Li, Zhiwen, et al.
Published: (2025)
by: Li, Zhiwen, et al.
Published: (2025)
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
by: Yu, Jiaao, et al.
Published: (2025)
by: Yu, Jiaao, et al.
Published: (2025)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
by: Ghazanfari, Sara, et al.
Published: (2025)
by: Ghazanfari, Sara, et al.
Published: (2025)
Object-Aware Video Matting with Cross-Frame Guidance
by: Zhang, Huayu, et al.
Published: (2025)
by: Zhang, Huayu, et al.
Published: (2025)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
by: Liu, Weijia, et al.
Published: (2025)
by: Liu, Weijia, et al.
Published: (2025)
A Spatio-Temporal based Frame Indexing Algorithm for QoS Improvement in Live Low-Motion Video Streaming
by: Adedokun, Adewale Emmanuel, et al.
Published: (2024)
by: Adedokun, Adewale Emmanuel, et al.
Published: (2024)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025)
by: Zhao, Zixu, et al.
Published: (2025)
Cost-Aware Routing for Efficient Text-To-Image Generation
by: Li, Qinchan, et al.
Published: (2025)
by: Li, Qinchan, et al.
Published: (2025)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025)
by: Huang, Jiehui, et al.
Published: (2025)
LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning
by: Chao, Lianying, et al.
Published: (2026)
by: Chao, Lianying, et al.
Published: (2026)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
LADDER: An Efficient Framework for Video Frame Interpolation
by: Shen, Tong, et al.
Published: (2024)
by: Shen, Tong, et al.
Published: (2024)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
by: Shen, Leqi, et al.
Published: (2024)
by: Shen, Leqi, et al.
Published: (2024)
Similar Items
-
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
by: Zhao, Zecheng, et al.
Published: (2026) -
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos
by: Zhao, Zecheng, et al.
Published: (2025) -
FastEdit: Fast Text-Guided Single-Image Editing via Semantic-Aware Diffusion Fine-Tuning
by: Chen, Zhi, et al.
Published: (2024) -
SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
by: Chen, Zhi, et al.
Published: (2025) -
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
by: Zhang, Deyu, et al.
Published: (2025)