TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
Fuente:
arXiv
Saved in:
| Main Authors: | Mishra, Pritam, Ballester, Coloma, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
by: Gomez, Lluis, et al.
Published: (2014)
by: Gomez, Lluis, et al.
Published: (2014)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
by: Díaz-Juan, Artur, et al.
Published: (2025)
by: Díaz-Juan, Artur, et al.
Published: (2025)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
Reading Between the Lanes: Text VideoQA on the Road
by: Tom, George, et al.
Published: (2023)
by: Tom, George, et al.
Published: (2023)
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025)
by: Ortega, Marc Serra, et al.
Published: (2025)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
by: Li, Zongmin, et al.
Published: (2026)
by: Li, Zongmin, et al.
Published: (2026)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026)
by: Fu, Xuanshuo, et al.
Published: (2026)
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024)
by: Llabrés, Artemis, et al.
Published: (2024)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
One missing piece in Vision and Language: A Survey on Comics Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
by: Sugihara, Tomoya, et al.
Published: (2024)
by: Sugihara, Tomoya, et al.
Published: (2024)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
by: Sarkar, Pritam, et al.
Published: (2025)
by: Sarkar, Pritam, et al.
Published: (2025)
ComicsPAP: understanding comic strips by picking the correct panel
by: Vivoli, Emanuele, et al.
Published: (2025)
by: Vivoli, Emanuele, et al.
Published: (2025)
Information-Maximized Soft Variable Discretization for Self-Supervised Image Representation Learning
by: Niu, Chuang, et al.
Published: (2025)
by: Niu, Chuang, et al.
Published: (2025)
Explicit Mutual Information Maximization for Self-Supervised Learning
by: Chang, Lele, et al.
Published: (2024)
by: Chang, Lele, et al.
Published: (2024)
Contrastive Learning under Noisy Temporal Self-Supervision for Colonoscopy Videos
by: Parolari, Luca, et al.
Published: (2026)
by: Parolari, Luca, et al.
Published: (2026)
VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
by: Mishra, Sandeep, et al.
Published: (2025)
by: Mishra, Sandeep, et al.
Published: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial Trimming
by: Yin, Zeyuan, et al.
Published: (2025)
by: Yin, Zeyuan, et al.
Published: (2025)
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
by: wang, Shuo, et al.
Published: (2025)
by: wang, Shuo, et al.
Published: (2025)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
Machine Unlearning for Document Classification
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
by: Xu, Zhen, et al.
Published: (2024)
by: Xu, Zhen, et al.
Published: (2024)
Leveraging Motion Information for Better Self-Supervised Video Correspondence Learning
by: Zhou, Zihan, et al.
Published: (2025)
by: Zhou, Zihan, et al.
Published: (2025)
RETHINED: A New Benchmark and Baseline for Real-Time High-Resolution Image Inpainting On Edge Devices
by: Sanchez, Marcelo, et al.
Published: (2025)
by: Sanchez, Marcelo, et al.
Published: (2025)
Information Maximization for Long-Tailed Semi-Supervised Domain Generalization
by: Fillioux, Leo, et al.
Published: (2026)
by: Fillioux, Leo, et al.
Published: (2026)
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
SelfHVD: Self-Supervised Handheld Video Deblurring
by: Xu, Honglei, et al.
Published: (2025)
by: Xu, Honglei, et al.
Published: (2025)
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
by: Sarkar, Pritam, et al.
Published: (2025)
by: Sarkar, Pritam, et al.
Published: (2025)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025)
by: Zuo, Zhi, et al.
Published: (2025)
Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
by: Sharma, Yash Kumar, et al.
Published: (2025)
by: Sharma, Yash Kumar, et al.
Published: (2025)
Similar Items
-
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026) -
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
by: Gomez, Lluis, et al.
Published: (2014) -
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
by: Díaz-Juan, Artur, et al.
Published: (2025) -
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024) -
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)