Exploring Efficient Foundational Multi-modal Models for Video Summarization
Fuente:
arXiv
Guardado en:
| Autores principales: | Samel, Karan, Beedu, Apoorva, Sontakke, Nitish, Essa, Irfan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
por: Samel, Karan, et al.
Publicado: (2025)
por: Samel, Karan, et al.
Publicado: (2025)
HierSum: A Global and Local Attention Mechanism for Video Summarization
por: Beedu, Apoorva, et al.
Publicado: (2025)
por: Beedu, Apoorva, et al.
Publicado: (2025)
On the Efficacy of Text-Based Input Modalities for Action Anticipation
por: Beedu, Apoorva, et al.
Publicado: (2024)
por: Beedu, Apoorva, et al.
Publicado: (2024)
Mamba Fusion: Learning Actions Through Questioning
por: Dong, Zhikang, et al.
Publicado: (2024)
por: Dong, Zhikang, et al.
Publicado: (2024)
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
por: Haresamudram, Harish, et al.
Publicado: (2024)
por: Haresamudram, Harish, et al.
Publicado: (2024)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
por: Marmon, Andrew, et al.
Publicado: (2024)
por: Marmon, Andrew, et al.
Publicado: (2024)
ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts
por: Choi, Sangbum, et al.
Publicado: (2025)
por: Choi, Sangbum, et al.
Publicado: (2025)
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data
por: Si, Haozhe, et al.
Publicado: (2025)
por: Si, Haozhe, et al.
Publicado: (2025)
An Integrated Framework for Multi-Granular Explanation of Video Summarization
por: Tsigos, Konstantinos, et al.
Publicado: (2024)
por: Tsigos, Konstantinos, et al.
Publicado: (2024)
Benchmarking Foundation Models for Zero-Shot Biometric Tasks
por: Sony, Redwan, et al.
Publicado: (2025)
por: Sony, Redwan, et al.
Publicado: (2025)
Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model
por: Delgrange, Camille, et al.
Publicado: (2024)
por: Delgrange, Camille, et al.
Publicado: (2024)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
por: He, Yuting, et al.
Publicado: (2026)
por: He, Yuting, et al.
Publicado: (2026)
Personalized Video Summarization by Multimodal Video Understanding
por: Chen, Brian, et al.
Publicado: (2024)
por: Chen, Brian, et al.
Publicado: (2024)
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
por: Khan, Anas Anwarul Haq, et al.
Publicado: (2025)
por: Khan, Anas Anwarul Haq, et al.
Publicado: (2025)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
por: Liang, Yujia, et al.
Publicado: (2025)
por: Liang, Yujia, et al.
Publicado: (2025)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
por: Park, Jungin, et al.
Publicado: (2025)
por: Park, Jungin, et al.
Publicado: (2025)
VideoSAGE: Video Summarization with Graph Representation Learning
por: Chaves, Jose M. Rojas, et al.
Publicado: (2024)
por: Chaves, Jose M. Rojas, et al.
Publicado: (2024)
Enhancing Video Summarization with Context Awareness
por: Huynh-Lam, Hai-Dang, et al.
Publicado: (2024)
por: Huynh-Lam, Hai-Dang, et al.
Publicado: (2024)
Comparing Learning Paradigms for Egocentric Video Summarization
por: Wen, Daniel
Publicado: (2025)
por: Wen, Daniel
Publicado: (2025)
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
por: Zhang, Zhiwang, et al.
Publicado: (2025)
por: Zhang, Zhiwang, et al.
Publicado: (2025)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
por: Yuan, Kun, et al.
Publicado: (2023)
por: Yuan, Kun, et al.
Publicado: (2023)
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
por: Mou, Tingshu, et al.
Publicado: (2026)
por: Mou, Tingshu, et al.
Publicado: (2026)
Cluster-based Video Summarization with Temporal Context Awareness
por: Huynh-Lam, Hai-Dang, et al.
Publicado: (2024)
por: Huynh-Lam, Hai-Dang, et al.
Publicado: (2024)
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
por: Khalil, Ahmad, et al.
Publicado: (2025)
por: Khalil, Ahmad, et al.
Publicado: (2025)
Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation
por: Zhou, Zhongliang, et al.
Publicado: (2024)
por: Zhou, Zhongliang, et al.
Publicado: (2024)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
por: Wang, Xin, et al.
Publicado: (2024)
por: Wang, Xin, et al.
Publicado: (2024)
Dense Video Captioning using Graph-based Sentence Summarization
por: Zhang, Zhiwang, et al.
Publicado: (2025)
por: Zhang, Zhiwang, et al.
Publicado: (2025)
An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
por: Eleftheriadis, Thomas, et al.
Publicado: (2025)
por: Eleftheriadis, Thomas, et al.
Publicado: (2025)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
por: Rajpal, Shreya, et al.
Publicado: (2026)
por: Rajpal, Shreya, et al.
Publicado: (2026)
VIMI: Grounding Video Generation through Multi-modal Instruction
por: Fang, Yuwei, et al.
Publicado: (2024)
por: Fang, Yuwei, et al.
Publicado: (2024)
Exploring Multi-modal Neural Scene Representations With Applications on Thermal Imaging
por: Özer, Mert, et al.
Publicado: (2024)
por: Özer, Mert, et al.
Publicado: (2024)
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
por: Liu, Shenxi, et al.
Publicado: (2025)
por: Liu, Shenxi, et al.
Publicado: (2025)
EdgeVidSum: Real-Time Personalized Video Summarization at the Edge
por: Mujtaba, Ghulam, et al.
Publicado: (2025)
por: Mujtaba, Ghulam, et al.
Publicado: (2025)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
por: Wu, Yuanli, et al.
Publicado: (2025)
por: Wu, Yuanli, et al.
Publicado: (2025)
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
por: Liu, Zhihang, et al.
Publicado: (2025)
por: Liu, Zhihang, et al.
Publicado: (2025)
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models through Reinforcement Learning from Ranking Feedback
por: Shi, Derek, et al.
Publicado: (2025)
por: Shi, Derek, et al.
Publicado: (2025)
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm
por: Goyal, Karan
Publicado: (2026)
por: Goyal, Karan
Publicado: (2026)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
por: Walimbe, Soham, et al.
Publicado: (2025)
por: Walimbe, Soham, et al.
Publicado: (2025)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
por: Liao, Pan, et al.
Publicado: (2026)
por: Liao, Pan, et al.
Publicado: (2026)
Multi-modal Auto-regressive Modeling via Visual Words
por: Peng, Tianshuo, et al.
Publicado: (2024)
por: Peng, Tianshuo, et al.
Publicado: (2024)
Ejemplares similares
-
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
por: Samel, Karan, et al.
Publicado: (2025) -
HierSum: A Global and Local Attention Mechanism for Video Summarization
por: Beedu, Apoorva, et al.
Publicado: (2025) -
On the Efficacy of Text-Based Input Modalities for Action Anticipation
por: Beedu, Apoorva, et al.
Publicado: (2024) -
Mamba Fusion: Learning Actions Through Questioning
por: Dong, Zhikang, et al.
Publicado: (2024) -
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
por: Haresamudram, Harish, et al.
Publicado: (2024)