Unsupervised Modality-Transferable Video Highlight Detection with Representation Activation Sequence Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Tingtian, Sun, Zixun, Xiao, Xinyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
Unsupervised Video Highlight Detection by Learning from Audio and Visual Recurrence
by: Islam, Zahidul, et al.
Published: (2024)
by: Islam, Zahidul, et al.
Published: (2024)
On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
by: Song, Xiufeng, et al.
Published: (2024)
by: Song, Xiufeng, et al.
Published: (2024)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
by: Barbakos, Spyros, et al.
Published: (2025)
by: Barbakos, Spyros, et al.
Published: (2025)
Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
Test-Time Adaptation for Video Highlight Detection Using Meta-Auxiliary Learning and Cross-Modality Hallucinations
by: Islam, Zahidul, et al.
Published: (2025)
by: Islam, Zahidul, et al.
Published: (2025)
HighlightMe: Detecting Highlights from Human-Centric Videos
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
Learning Natural Consistency Representation for Face Forgery Video Detection
by: Zhang, Daichi, et al.
Published: (2024)
by: Zhang, Daichi, et al.
Published: (2024)
Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification
by: Yang, Xiaomei, et al.
Published: (2026)
by: Yang, Xiaomei, et al.
Published: (2026)
CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
by: Dong, Xin, et al.
Published: (2026)
by: Dong, Xin, et al.
Published: (2026)
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
Let Video Teaches You More: Video-to-Image Knowledge Distillation using DEtection TRansformer for Medical Video Lesion Detection
by: Jiang, Yuncheng, et al.
Published: (2024)
by: Jiang, Yuncheng, et al.
Published: (2024)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
by: Zhao, Henghao, et al.
Published: (2023)
by: Zhao, Henghao, et al.
Published: (2023)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
by: Ma, Wenxuan, et al.
Published: (2024)
by: Ma, Wenxuan, et al.
Published: (2024)
Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
by: Cai, Lincan, et al.
Published: (2024)
by: Cai, Lincan, et al.
Published: (2024)
Cross-Modal Prototype Allocation: Unsupervised Slide Representation Learning via Patch-Text Contrast in Computational Pathology
by: Chen, Yuxuan, et al.
Published: (2025)
by: Chen, Yuxuan, et al.
Published: (2025)
Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection
by: He, Liren, et al.
Published: (2024)
by: He, Liren, et al.
Published: (2024)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
Unleash the Potential of CLIP for Video Highlight Detection
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
Exploring Diffusion Time-steps for Unsupervised Representation Learning
by: Yue, Zhongqi, et al.
Published: (2024)
by: Yue, Zhongqi, et al.
Published: (2024)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
by: Jiang, Lei, et al.
Published: (2025)
by: Jiang, Lei, et al.
Published: (2025)
Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection
by: Yang, Hui-Yue, et al.
Published: (2024)
by: Yang, Hui-Yue, et al.
Published: (2024)
MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning
by: Ma, Hongxu, et al.
Published: (2025)
by: Ma, Hongxu, et al.
Published: (2025)
No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D Consistency
by: Wu, Cho-Ying, et al.
Published: (2026)
by: Wu, Cho-Ying, et al.
Published: (2026)
CheXLearner: Text-Guided Fine-Grained Representation Learning for Progression Detection
by: Wang, Yuanzhuo, et al.
Published: (2025)
by: Wang, Yuanzhuo, et al.
Published: (2025)
GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation
by: Li, Chenxin, et al.
Published: (2024)
by: Li, Chenxin, et al.
Published: (2024)
Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection
by: Shao, YiKang, et al.
Published: (2025)
by: Shao, YiKang, et al.
Published: (2025)
Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies
by: Tian, Mulin, et al.
Published: (2023)
by: Tian, Mulin, et al.
Published: (2023)
CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance Videos
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Unsupervised Learning of Disentangled Representations from Video
by: Denton, Remi, et al.
Published: (2017)
by: Denton, Remi, et al.
Published: (2017)
On the Transferability and Discriminability of Repersentation Learning in Unsupervised Domain Adaptation
by: Qiang, Wenwen, et al.
Published: (2025)
by: Qiang, Wenwen, et al.
Published: (2025)
Unified Unsupervised Salient Object Detection via Knowledge Transfer
by: Yuan, Yao, et al.
Published: (2024)
by: Yuan, Yao, et al.
Published: (2024)
Jenga Stacking Based on 6D Pose Estimation for Architectural Form Finding Process
by: Huang, Zixun
Published: (2023)
by: Huang, Zixun
Published: (2023)
Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking
by: Chen, Xin, et al.
Published: (2023)
by: Chen, Xin, et al.
Published: (2023)
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
by: Zhang, Yuechen, et al.
Published: (2023)
by: Zhang, Yuechen, et al.
Published: (2023)
Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection
by: Li, Jinglun, et al.
Published: (2024)
by: Li, Jinglun, et al.
Published: (2024)
Semantic Feature Learning for Universal Unsupervised Cross-Domain Retrieval
by: Wang, Lixu, et al.
Published: (2024)
by: Wang, Lixu, et al.
Published: (2024)
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
by: Li, Xinyao, et al.
Published: (2024)
by: Li, Xinyao, et al.
Published: (2024)
Learning Transferable Negative Prompts for Out-of-Distribution Detection
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
Contrast-Phys: Unsupervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
by: Sun, Zhaodong, et al.
Published: (2022)
by: Sun, Zhaodong, et al.
Published: (2022)
Similar Items
-
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026) -
Unsupervised Video Highlight Detection by Learning from Audio and Visual Recurrence
by: Islam, Zahidul, et al.
Published: (2024) -
On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
by: Song, Xiufeng, et al.
Published: (2024) -
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
by: Barbakos, Spyros, et al.
Published: (2025) -
Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
by: Yang, Dingkang, et al.
Published: (2024)