GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yunzhuo, Xu, Yifang, Xie, Zien, Shu, Yukun, Du, Sidan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
by: Xu, Yifang, et al.
Published: (2024)
by: Xu, Yifang, et al.
Published: (2024)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
Pyramid Feature Attention Network for Monocular Depth Prediction
by: Xu, Yifang, et al.
Published: (2024)
by: Xu, Yifang, et al.
Published: (2024)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
by: Paul, Dhiman, et al.
Published: (2024)
by: Paul, Dhiman, et al.
Published: (2024)
FaceSnap: Enhanced ID-fidelity Network for Tuning-free Portrait Customization
by: Zhai, Benxiang, et al.
Published: (2026)
by: Zhai, Benxiang, et al.
Published: (2026)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
by: Yang, Jin, et al.
Published: (2024)
by: Yang, Jin, et al.
Published: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
by: Sun, Yunzhuo, et al.
Published: (2026)
by: Sun, Yunzhuo, et al.
Published: (2026)
Saliency-Guided DETR for Moment Retrieval and Highlight Detection
by: Gordeev, Aleksandr, et al.
Published: (2024)
by: Gordeev, Aleksandr, et al.
Published: (2024)
Animating the Past: Reconstruct Trilobite via Video Generation
by: Wu, Xiaoran, et al.
Published: (2024)
by: Wu, Xiaoran, et al.
Published: (2024)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
by: Cao, Zhuo, et al.
Published: (2025)
by: Cao, Zhuo, et al.
Published: (2025)
Moment and Highlight Detection via MLLM Frame Segmentation
by: Jiwanta, I Putu Andika Bagas, et al.
Published: (2025)
by: Jiwanta, I Putu Andika Bagas, et al.
Published: (2025)
Proto-OOD: Enhancing OOD Object Detection with Prototype Feature Similarity
by: Chen, Junkun, et al.
Published: (2024)
by: Chen, Junkun, et al.
Published: (2024)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
by: Zhao, Henghao, et al.
Published: (2023)
by: Zhao, Henghao, et al.
Published: (2023)
Diff-PC: Identity-preserving and 3D-aware Controllable Diffusion for Zero-shot Portrait Customization
by: Xu, Yifang, et al.
Published: (2026)
by: Xu, Yifang, et al.
Published: (2026)
SemanticMoments: Training-Free Motion Similarity via Third Moment Features
by: Huberman, Saar, et al.
Published: (2026)
by: Huberman, Saar, et al.
Published: (2026)
CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
by: Dong, Xin, et al.
Published: (2026)
by: Dong, Xin, et al.
Published: (2026)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
by: Lee, YuEun, et al.
Published: (2025)
by: Lee, YuEun, et al.
Published: (2025)
MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning
by: Ma, Hongxu, et al.
Published: (2025)
by: Ma, Hongxu, et al.
Published: (2025)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
by: Yu, An, et al.
Published: (2025)
by: Yu, An, et al.
Published: (2025)
DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
by: Li, Binbin, et al.
Published: (2025)
by: Li, Binbin, et al.
Published: (2025)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
by: Park, Seojeong, et al.
Published: (2024)
by: Park, Seojeong, et al.
Published: (2024)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Feature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detection
by: Zhao, Rongqiang, et al.
Published: (2026)
by: Zhao, Rongqiang, et al.
Published: (2026)
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Unleash the Potential of CLIP for Video Highlight Detection
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
Reference-Free Omnidirectional Stereo Matching via Multi-View Consistency Maximization
by: Xu, Lehuai, et al.
Published: (2026)
by: Xu, Lehuai, et al.
Published: (2026)
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
by: Jiang, Chen, et al.
Published: (2023)
by: Jiang, Chen, et al.
Published: (2023)
Dual-Mode Deep Anomaly Detection for Medical Manufacturing: Structural Similarity and Feature Distance
by: Diaz, Julio Zanon, et al.
Published: (2025)
by: Diaz, Julio Zanon, et al.
Published: (2025)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
TS-SatFire: A Multi-Task Satellite Image Time-Series Dataset for Wildfire Detection and Prediction
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
by: Zhao, Pengcheng, et al.
Published: (2025)
by: Zhao, Pengcheng, et al.
Published: (2025)
SLVideo: A Sign Language Video Moment Retrieval Framework
by: Martins, Gonçalo Vinagre, et al.
Published: (2024)
by: Martins, Gonçalo Vinagre, et al.
Published: (2024)
Deep Learning-Based Fatigue Cracks Detection in Bridge Girders using Feature Pyramid Networks
by: Zhang, Jiawei, et al.
Published: (2024)
by: Zhang, Jiawei, et al.
Published: (2024)
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
by: Zhang, Shihang, et al.
Published: (2026)
by: Zhang, Shihang, et al.
Published: (2026)
Efficient CNN Compression via Multi-method Low Rank Factorization and Feature Map Similarity
by: Kokhazadeh, M., et al.
Published: (2025)
by: Kokhazadeh, M., et al.
Published: (2025)
Similar Items
-
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
by: Xu, Yifang, et al.
Published: (2025) -
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
by: Xu, Yifang, et al.
Published: (2024) -
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025) -
HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
by: Xu, Yifang, et al.
Published: (2025) -
Pyramid Feature Attention Network for Monocular Depth Prediction
by: Xu, Yifang, et al.
Published: (2024)