A Flexible and Scalable Framework for Video Moment Search
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chongzhi, Zhu, Xizhou, Sun, Aixin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Event-aware Video Corpus Moment Retrieval
von: Hou, Danyang, et al.
Veröffentlicht: (2024)
von: Hou, Danyang, et al.
Veröffentlicht: (2024)
Improving Video Corpus Moment Retrieval with Partial Relevance Enhancement
von: Hou, Danyang, et al.
Veröffentlicht: (2024)
von: Hou, Danyang, et al.
Veröffentlicht: (2024)
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
von: Tang, Hengzhu, et al.
Veröffentlicht: (2025)
von: Tang, Hengzhu, et al.
Veröffentlicht: (2025)
LongVidSearch: An Agentic Benchmark for Multi-hop Evidence Retrieval Planning in Long Videos
von: Yu, Rongyi, et al.
Veröffentlicht: (2026)
von: Yu, Rongyi, et al.
Veröffentlicht: (2026)
Video Editing for Video Retrieval
von: Zhu, Bin, et al.
Veröffentlicht: (2024)
von: Zhu, Bin, et al.
Veröffentlicht: (2024)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2025)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2025)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
Multi-scale 2D Temporal Map Diffusion Models for Natural Language Video Localization
von: Zhang, Chongzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Chongzhi, et al.
Veröffentlicht: (2024)
Scalable Residual Feature Aggregation Framework with Hybrid Metaheuristic Optimization for Robust Early Pancreatic Neoplasm Detection in Multimodal CT Imaging
von: Thiruvengadam, Janani Annur, et al.
Veröffentlicht: (2025)
von: Thiruvengadam, Janani Annur, et al.
Veröffentlicht: (2025)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
von: Hu, Fan, et al.
Veröffentlicht: (2025)
von: Hu, Fan, et al.
Veröffentlicht: (2025)
FF-PNet: A Pyramid Network Based on Feature and Field for Brain Image Registration
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
NextAds: Towards Next-generation Personalized Video Advertising
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
A Novel Evaluation Framework for Image2Text Generation
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Visual Product Search Benchmark
von: Govindappa, Karthik Sulthanpete
Veröffentlicht: (2026)
von: Govindappa, Karthik Sulthanpete
Veröffentlicht: (2026)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation
von: Chakraborty, Debashish, et al.
Veröffentlicht: (2026)
von: Chakraborty, Debashish, et al.
Veröffentlicht: (2026)
Diff-VPS: Video Polyp Segmentation via a Multi-task Diffusion Network with Adversarial Temporal Reasoning
von: Lu, Yingling, et al.
Veröffentlicht: (2024)
von: Lu, Yingling, et al.
Veröffentlicht: (2024)
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
von: Reddy, Arun, et al.
Veröffentlicht: (2025)
von: Reddy, Arun, et al.
Veröffentlicht: (2025)
Smart Routing for Multimodal Video Retrieval: When to Search What
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
YOLO-Vehicle-Pro: A Cloud-Edge Collaborative Framework for Object Detection in Autonomous Driving under Adverse Weather Conditions
von: Li, Xiguang, et al.
Veröffentlicht: (2024)
von: Li, Xiguang, et al.
Veröffentlicht: (2024)
Adapting MLLMs for Nuanced Video Retrieval
von: Bagad, Piyush, et al.
Veröffentlicht: (2025)
von: Bagad, Piyush, et al.
Veröffentlicht: (2025)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
Interactive Mars Image Content-Based Search with Interpretable Machine Learning
von: Vasu, Bhavan, et al.
Veröffentlicht: (2024)
von: Vasu, Bhavan, et al.
Veröffentlicht: (2024)
TrajSV: A Trajectory-based Model for Sports Video Representations and Applications
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
von: Skow, Tyler, et al.
Veröffentlicht: (2026)
von: Skow, Tyler, et al.
Veröffentlicht: (2026)
DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
When & How to Write for Personalized Demand-aware Query Rewriting in Video Search
von: cheng, Cheng, et al.
Veröffentlicht: (2025)
von: cheng, Cheng, et al.
Veröffentlicht: (2025)
Heterogeneous Graph-based Framework with Disentangled Representations Learning for Multi-target Cross Domain Recommendation
von: Liu, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Liu, Xiaopeng, et al.
Veröffentlicht: (2024)
Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
von: Lin, Junan, et al.
Veröffentlicht: (2025)
von: Lin, Junan, et al.
Veröffentlicht: (2025)
Advancing Re-Ranking with Multimodal Fusion and Target-Oriented Auxiliary Tasks in E-Commerce Search
von: Xu, Enqiang, et al.
Veröffentlicht: (2024)
von: Xu, Enqiang, et al.
Veröffentlicht: (2024)
Multimodal Language Models for Domain-Specific Procedural Video Summarization
von: Hussain, Nafisa
Veröffentlicht: (2024)
von: Hussain, Nafisa
Veröffentlicht: (2024)
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
A Multi-Granularity Retrieval Framework for Visually-Rich Documents
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
von: Li, Minghan, et al.
Veröffentlicht: (2026)
von: Li, Minghan, et al.
Veröffentlicht: (2026)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Event-aware Video Corpus Moment Retrieval
von: Hou, Danyang, et al.
Veröffentlicht: (2024) -
Improving Video Corpus Moment Retrieval with Partial Relevance Enhancement
von: Hou, Danyang, et al.
Veröffentlicht: (2024) -
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
von: Tang, Hengzhu, et al.
Veröffentlicht: (2025) -
LongVidSearch: An Agentic Benchmark for Multi-hop Evidence Retrieval Planning in Long Videos
von: Yu, Rongyi, et al.
Veröffentlicht: (2026) -
Video Editing for Video Retrieval
von: Zhu, Bin, et al.
Veröffentlicht: (2024)