TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Mishra, Pritam, Ballester, Coloma, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025)
by: Mishra, Pritam, et al.
Published: (2025)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
by: Díaz-Juan, Artur, et al.
Published: (2025)
by: Díaz-Juan, Artur, et al.
Published: (2025)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
Comparing Learning Paradigms for Egocentric Video Summarization
by: Wen, Daniel
Published: (2025)
by: Wen, Daniel
Published: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
Machine Unlearning for Document Classification
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
by: Singh, Azad, et al.
Published: (2025)
by: Singh, Azad, et al.
Published: (2025)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
by: Gomez, Lluis, et al.
Published: (2014)
by: Gomez, Lluis, et al.
Published: (2014)
Self-Supervised Learning for Endoscopic Video Analysis
by: Hirsch, Roy, et al.
Published: (2023)
by: Hirsch, Roy, et al.
Published: (2023)
VideoSAGE: Video Summarization with Graph Representation Learning
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Rethinking Deep Clustering Paradigms: Self-Supervision Is All You Need
by: Shaheena, Amal, et al.
Published: (2025)
by: Shaheena, Amal, et al.
Published: (2025)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
by: Abbasi, Mehryar, et al.
Published: (2024)
by: Abbasi, Mehryar, et al.
Published: (2024)
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
by: Liu, Yuhong, et al.
Published: (2025)
by: Liu, Yuhong, et al.
Published: (2025)
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
MAEPose: Self-Supervised Spatiotemporal Learning for Human Pose Estimation on mmWave Video
by: Wei, Xijia, et al.
Published: (2026)
by: Wei, Xijia, et al.
Published: (2026)
Image-Feature Weak-to-Strong Consistency: An Enhanced Paradigm for Semi-Supervised Learning
by: Wu, Zhiyu, et al.
Published: (2024)
by: Wu, Zhiyu, et al.
Published: (2024)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
by: Yue, Feng, et al.
Published: (2025)
by: Yue, Feng, et al.
Published: (2025)
An Integrated Framework for Multi-Granular Explanation of Video Summarization
by: Tsigos, Konstantinos, et al.
Published: (2024)
by: Tsigos, Konstantinos, et al.
Published: (2024)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange
by: Wu, Yanhao, et al.
Published: (2024)
by: Wu, Yanhao, et al.
Published: (2024)
ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding
by: Chen, Kehua
Published: (2025)
by: Chen, Kehua
Published: (2025)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
by: Rajpal, Shreya, et al.
Published: (2026)
by: Rajpal, Shreya, et al.
Published: (2026)
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
by: Eleftheriadis, Thomas, et al.
Published: (2025)
by: Eleftheriadis, Thomas, et al.
Published: (2025)
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024)
by: Samel, Karan, et al.
Published: (2024)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
by: Han, Ruiyan, et al.
Published: (2026)
by: Han, Ruiyan, et al.
Published: (2026)
Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
by: Sun, Shengyang, et al.
Published: (2026)
by: Sun, Shengyang, et al.
Published: (2026)
EdgeVidSum: Real-Time Personalized Video Summarization at the Edge
by: Mujtaba, Ghulam, et al.
Published: (2025)
by: Mujtaba, Ghulam, et al.
Published: (2025)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
by: Wu, Yuanli, et al.
Published: (2025)
by: Wu, Yuanli, et al.
Published: (2025)
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
Enhancing Self-Supervised Fine-Grained Video Object Tracking with Dynamic Memory Prediction
by: Zhou, Zihan, et al.
Published: (2025)
by: Zhou, Zihan, et al.
Published: (2025)
Automatized Self-Supervised Learning for Skin Lesion Screening
by: Useini, Vullnet, et al.
Published: (2023)
by: Useini, Vullnet, et al.
Published: (2023)
DINO-MX: A Modular & Flexible Framework for Self-Supervised Learning
by: Gokmen, Mahmut Selman, et al.
Published: (2025)
by: Gokmen, Mahmut Selman, et al.
Published: (2025)
Similar Items
-
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025) -
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
by: Díaz-Juan, Artur, et al.
Published: (2025) -
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024) -
Comparing Learning Paradigms for Egocentric Video Summarization
by: Wen, Daniel
Published: (2025) -
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)