Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
Fuente:
arXiv
Saved in:
| Main Authors: | Sugihara, Tomoya, Masuda, Shuntaro, Xiao, Ling, Yamasaki, Toshihiko |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
by: Xiao, Ling, et al.
Published: (2022)
by: Xiao, Ling, et al.
Published: (2022)
Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
by: Tanji, Naoto, et al.
Published: (2025)
by: Tanji, Naoto, et al.
Published: (2025)
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
by: Xiao, Ling, et al.
Published: (2026)
by: Xiao, Ling, et al.
Published: (2026)
Language-Guided Graph Representation Learning for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
by: Barbara, Mario, et al.
Published: (2025)
by: Barbara, Mario, et al.
Published: (2025)
Language-guided Detection and Mitigation of Unknown Dataset Bias
by: Zhao, Zaiying, et al.
Published: (2024)
by: Zhao, Zaiying, et al.
Published: (2024)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
Face2Diffusion for Fast and Editable Face Personalization
by: Shiohara, Kaede, et al.
Published: (2024)
by: Shiohara, Kaede, et al.
Published: (2024)
Unified Vector Floorplan Generation via Markup Representation
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
by: Kuroki, Michihiro, et al.
Published: (2025)
by: Kuroki, Michihiro, et al.
Published: (2025)
BSED: Baseline Shapley-Based Explainable Detector
by: Kuroki, Michihiro, et al.
Published: (2023)
by: Kuroki, Michihiro, et al.
Published: (2023)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
by: Zhao, Zaiying, et al.
Published: (2025)
by: Zhao, Zaiying, et al.
Published: (2025)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
by: Mu, Kuan-Chen, et al.
Published: (2024)
by: Mu, Kuan-Chen, et al.
Published: (2024)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025)
by: Mishra, Pritam, et al.
Published: (2025)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
Online Open-set Semi-supervised Object Detection with Dual Competing Head
by: Wang, Zerun, et al.
Published: (2023)
by: Wang, Zerun, et al.
Published: (2023)
Spectral Probing of Feature Upsamplers in 2D-to-3D Scene Reconstruction
by: Xiao, Ling, et al.
Published: (2026)
by: Xiao, Ling, et al.
Published: (2026)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
by: Chen, Chao, et al.
Published: (2023)
by: Chen, Chao, et al.
Published: (2023)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024)
by: Kulkarni, Yogesh, et al.
Published: (2024)
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection
by: Lee, Junhee, et al.
Published: (2025)
by: Lee, Junhee, et al.
Published: (2025)
Reward Incremental Learning in Text-to-Image Generation
by: Wang, Maorong, et al.
Published: (2024)
by: Wang, Maorong, et al.
Published: (2024)
SCOMatch: Alleviating Overtrusting in Open-set Semi-supervised Learning
by: Wang, Zerun, et al.
Published: (2024)
by: Wang, Zerun, et al.
Published: (2024)
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
by: Sbrolli, Cristian, et al.
Published: (2026)
by: Sbrolli, Cristian, et al.
Published: (2026)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation
by: Xing, Yun, et al.
Published: (2023)
by: Xing, Yun, et al.
Published: (2023)
ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing Points
by: Okumura, Ryota, et al.
Published: (2025)
by: Okumura, Ryota, et al.
Published: (2025)
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
by: Ahmadian, Mona, et al.
Published: (2024)
by: Ahmadian, Mona, et al.
Published: (2024)
Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
Text-Guided Video Masked Autoencoder
by: Fan, David, et al.
Published: (2024)
by: Fan, David, et al.
Published: (2024)
VideoClusterNet: Self-Supervised and Adaptive Face Clustering For Videos
by: Walawalkar, Devesh, et al.
Published: (2024)
by: Walawalkar, Devesh, et al.
Published: (2024)
SelfHVD: Self-Supervised Handheld Video Deblurring
by: Xu, Honglei, et al.
Published: (2025)
by: Xu, Honglei, et al.
Published: (2025)
Scaling Up Video Summarization Pretraining with Large Language Models
by: Argaw, Dawit Mureja, et al.
Published: (2024)
by: Argaw, Dawit Mureja, et al.
Published: (2024)
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
by: Shi, Miaojing, et al.
Published: (2026)
by: Shi, Miaojing, et al.
Published: (2026)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
by: Peng, Bo, et al.
Published: (2023)
by: Peng, Bo, et al.
Published: (2023)
Self-Supervised Video Desmoking for Laparoscopic Surgery
by: Wu, Renlong, et al.
Published: (2024)
by: Wu, Renlong, et al.
Published: (2024)
Self-Supervised Animal Identification for Long Videos
by: Fang, Xuyang, et al.
Published: (2026)
by: Fang, Xuyang, et al.
Published: (2026)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
by: Rajpal, Shreya, et al.
Published: (2026)
by: Rajpal, Shreya, et al.
Published: (2026)
Similar Items
-
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
by: Xiao, Ling, et al.
Published: (2022) -
Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
by: Tanji, Naoto, et al.
Published: (2025) -
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
by: Xiao, Ling, et al.
Published: (2026) -
Language-Guided Graph Representation Learning for Video Summarization
by: Li, Wenrui, et al.
Published: (2025) -
Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
by: Barbara, Mario, et al.
Published: (2025)