Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zecheng, Chen, Zhi, Huang, Zi, Sadiq, Shazia, Chen, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
by: Zhao, Zecheng, et al.
Published: (2025)
by: Zhao, Zecheng, et al.
Published: (2025)
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos
by: Zhao, Zecheng, et al.
Published: (2025)
by: Zhao, Zecheng, et al.
Published: (2025)
FastEdit: Fast Text-Guided Single-Image Editing via Semantic-Aware Diffusion Fine-Tuning
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
by: Skow, Tyler, et al.
Published: (2026)
by: Skow, Tyler, et al.
Published: (2026)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation
by: Yuan, Chao, et al.
Published: (2026)
by: Yuan, Chao, et al.
Published: (2026)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024)
by: Zhang, Haonan, et al.
Published: (2024)
MTLSI-Net: A Linear Semantic Interaction Network for Parameter-Efficient Multi-Task Dense Prediction
by: Liu, Chen, et al.
Published: (2026)
by: Liu, Chen, et al.
Published: (2026)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
A Dense Reward View on Aligning Text-to-Image Diffusion with Preference
by: Yang, Shentao, et al.
Published: (2024)
by: Yang, Shentao, et al.
Published: (2024)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
by: Tian, Kaibin, et al.
Published: (2024)
by: Tian, Kaibin, et al.
Published: (2024)
T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
by: Li, Yili, et al.
Published: (2024)
by: Li, Yili, et al.
Published: (2024)
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
by: Huang, Yibin, et al.
Published: (2025)
by: Huang, Yibin, et al.
Published: (2025)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
by: Luo, Dezhao, et al.
Published: (2024)
by: Luo, Dezhao, et al.
Published: (2024)
Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence
by: Filntisis, Panagiotis P., et al.
Published: (2026)
by: Filntisis, Panagiotis P., et al.
Published: (2026)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
by: Shen, Leqi, et al.
Published: (2024)
by: Shen, Leqi, et al.
Published: (2024)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025)
by: Zhao, Zixu, et al.
Published: (2025)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
by: Zhang, Deyu, et al.
Published: (2025)
by: Zhang, Deyu, et al.
Published: (2025)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
by: Chen, Weiliang, et al.
Published: (2025)
by: Chen, Weiliang, et al.
Published: (2025)
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
by: Wang, Xiaoye, et al.
Published: (2025)
by: Wang, Xiaoye, et al.
Published: (2025)
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
by: Li, Wenjun, et al.
Published: (2024)
by: Li, Wenjun, et al.
Published: (2024)
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking
by: Tran, Huu-Loc, et al.
Published: (2025)
by: Tran, Huu-Loc, et al.
Published: (2025)
From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
by: Li, Ling, et al.
Published: (2026)
by: Li, Ling, et al.
Published: (2026)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
DKPMV: Dense Keypoints Fusion from Multi-View RGB Frames for 6D Pose Estimation of Textureless Objects
by: Chen, Jiahong, et al.
Published: (2025)
by: Chen, Jiahong, et al.
Published: (2025)
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
by: Sun, Shuo, et al.
Published: (2026)
by: Sun, Shuo, et al.
Published: (2026)
SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language Retrieval
by: Jiang, Longtao, et al.
Published: (2024)
by: Jiang, Longtao, et al.
Published: (2024)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval
by: Kang, Bin, et al.
Published: (2025)
by: Kang, Bin, et al.
Published: (2025)
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
by: Yang, Yunhan, et al.
Published: (2025)
by: Yang, Yunhan, et al.
Published: (2025)
Adversarial Video Promotion Against Text-to-Video Retrieval
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
R3G: A Reasoning--Retrieval--Reranking Framework for Vision-Centric Answer Generation
by: Chen, Zhuohong, et al.
Published: (2026)
by: Chen, Zhuohong, et al.
Published: (2026)
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
by: Tang, Haoran, et al.
Published: (2024)
by: Tang, Haoran, et al.
Published: (2024)
Dense Video Captioning Using Unsupervised Semantic Information
by: Estevam, Valter, et al.
Published: (2021)
by: Estevam, Valter, et al.
Published: (2021)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router
by: Huang, Yubo, et al.
Published: (2025)
by: Huang, Yubo, et al.
Published: (2025)
TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
Similar Items
-
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
by: Zhao, Zecheng, et al.
Published: (2025) -
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos
by: Zhao, Zecheng, et al.
Published: (2025) -
FastEdit: Fast Text-Guided Single-Image Editing via Semantic-Aware Diffusion Fine-Tuning
by: Chen, Zhi, et al.
Published: (2024) -
SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
by: Chen, Zhi, et al.
Published: (2025) -
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
by: Skow, Tyler, et al.
Published: (2026)