Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Junkai, Wang, Qirui, Jin, Yaoqing, Ma, Shuai, Xu, Minghan, Pang, Shanmin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918533098635264
author Yang, Junkai
Wang, Qirui
Jin, Yaoqing
Ma, Shuai
Xu, Minghan
Pang, Shanmin
author_facet Yang, Junkai
Wang, Qirui
Jin, Yaoqing
Ma, Shuai
Xu, Minghan
Pang, Shanmin
contents Retrieving partially relevant segments from untrimmed videos remains difficult due to two persistent challenges: the mismatch in information density between text and video segments, and limited attention mechanisms that overlook semantic focus and event correlations. We present KDC-Net, a Knowledge-Refined Dual Context-Aware Network that tackles these issues from both textual and visual perspectives. On the text side, a Hierarchical Semantic Aggregation module captures and adaptively fuses multi-scale phrase cues to enrich query semantics. On the video side, a Dynamic Temporal Attention mechanism employs relative positional encoding and adaptive temporal windows to highlight key events with local temporal coherence. Additionally, a dynamic CLIP-based distillation strategy, enhanced with temporal-continuity-aware refinement, ensures segment-aware and objective-aligned knowledge transfer. Experiments on PRVR benchmarks show that KDC-Net consistently outperforms state-of-the-art methods, especially under low moment-to-video ratios.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23902
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval
Yang, Junkai
Wang, Qirui
Jin, Yaoqing
Ma, Shuai
Xu, Minghan
Pang, Shanmin
Computer Vision and Pattern Recognition
Artificial Intelligence
Retrieving partially relevant segments from untrimmed videos remains difficult due to two persistent challenges: the mismatch in information density between text and video segments, and limited attention mechanisms that overlook semantic focus and event correlations. We present KDC-Net, a Knowledge-Refined Dual Context-Aware Network that tackles these issues from both textual and visual perspectives. On the text side, a Hierarchical Semantic Aggregation module captures and adaptively fuses multi-scale phrase cues to enrich query semantics. On the video side, a Dynamic Temporal Attention mechanism employs relative positional encoding and adaptive temporal windows to highlight key events with local temporal coherence. Additionally, a dynamic CLIP-based distillation strategy, enhanced with temporal-continuity-aware refinement, ensures segment-aware and objective-aligned knowledge transfer. Experiments on PRVR benchmarks show that KDC-Net consistently outperforms state-of-the-art methods, especially under low moment-to-video ratios.
title Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.23902