Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
Fuente:
arXiv
Guardado en:
| Autores principales: | Samel, Karan, Sontakke, Nitish, Essa, Irfan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring Efficient Foundational Multi-modal Models for Video Summarization
por: Samel, Karan, et al.
Publicado: (2024)
por: Samel, Karan, et al.
Publicado: (2024)
On the Efficacy of Text-Based Input Modalities for Action Anticipation
por: Beedu, Apoorva, et al.
Publicado: (2024)
por: Beedu, Apoorva, et al.
Publicado: (2024)
HierSum: A Global and Local Attention Mechanism for Video Summarization
por: Beedu, Apoorva, et al.
Publicado: (2025)
por: Beedu, Apoorva, et al.
Publicado: (2025)
Efficient Pre-training for Localized Instruction Generation of Videos
por: Batra, Anil, et al.
Publicado: (2023)
por: Batra, Anil, et al.
Publicado: (2023)
SLAIM: Robust Dense Neural SLAM for Online Tracking and Mapping
por: Cartillier, Vincent, et al.
Publicado: (2024)
por: Cartillier, Vincent, et al.
Publicado: (2024)
3D Semantic MapNet: Building Maps for Multi-Object Re-Identification in 3D
por: Cartillier, Vincent, et al.
Publicado: (2024)
por: Cartillier, Vincent, et al.
Publicado: (2024)
Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
por: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Publicado: (2024)
por: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Publicado: (2024)
Open-Event Procedure Planning in Instructional Videos
por: Wu, Yilu, et al.
Publicado: (2024)
por: Wu, Yilu, et al.
Publicado: (2024)
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
por: Zhao, Jinghan, et al.
Publicado: (2025)
por: Zhao, Jinghan, et al.
Publicado: (2025)
ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
por: Seminara, Luigi, et al.
Publicado: (2026)
por: Seminara, Luigi, et al.
Publicado: (2026)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
por: Chen, Lan, et al.
Publicado: (2025)
por: Chen, Lan, et al.
Publicado: (2025)
RECIPE: Procedural Planning via Grounding in Instructional Video
por: Seminara, Luigi, et al.
Publicado: (2026)
por: Seminara, Luigi, et al.
Publicado: (2026)
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
por: Wang, Hanlin, et al.
Publicado: (2023)
por: Wang, Hanlin, et al.
Publicado: (2023)
Predicting Implicit Arguments in Procedural Video Instructions
por: Batra, Anil, et al.
Publicado: (2025)
por: Batra, Anil, et al.
Publicado: (2025)
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
por: Zhou, Yufan, et al.
Publicado: (2025)
por: Zhou, Yufan, et al.
Publicado: (2025)
SG-MIM: Structured Knowledge Guided Efficient Pre-training for Dense Prediction
por: Son, Sumin, et al.
Publicado: (2024)
por: Son, Sumin, et al.
Publicado: (2024)
MoCHA: Denoising Caption Supervision for Motion-Text Retrieval
por: Warner, Nikolai, et al.
Publicado: (2026)
por: Warner, Nikolai, et al.
Publicado: (2026)
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
por: Orlova, Svetlana, et al.
Publicado: (2026)
por: Orlova, Svetlana, et al.
Publicado: (2026)
VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
por: Zhang, Ziyang, et al.
Publicado: (2025)
por: Zhang, Ziyang, et al.
Publicado: (2025)
Contrastive Language Video Time Pre-training
por: Liu, Hengyue, et al.
Publicado: (2024)
por: Liu, Hengyue, et al.
Publicado: (2024)
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training
por: Tian, Qingyao, et al.
Publicado: (2025)
por: Tian, Qingyao, et al.
Publicado: (2025)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
por: Wu, Te-Lin, et al.
Publicado: (2021)
por: Wu, Te-Lin, et al.
Publicado: (2021)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
por: Marmon, Andrew, et al.
Publicado: (2024)
por: Marmon, Andrew, et al.
Publicado: (2024)
LAP: A Language-Aware Planning Model For Procedure Planning In Instructional Videos
por: Shi, Lei, et al.
Publicado: (2026)
por: Shi, Lei, et al.
Publicado: (2026)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
por: Shi, Lei, et al.
Publicado: (2024)
por: Shi, Lei, et al.
Publicado: (2024)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
por: Fernando, Basura, et al.
Publicado: (2025)
por: Fernando, Basura, et al.
Publicado: (2025)
Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
por: Dong, Yuhao, et al.
Publicado: (2026)
por: Dong, Yuhao, et al.
Publicado: (2026)
Mamba Fusion: Learning Actions Through Questioning
por: Dong, Zhikang, et al.
Publicado: (2024)
por: Dong, Zhikang, et al.
Publicado: (2024)
Leveraging Pre-trained CNNs for Efficient Feature Extraction in Rice Leaf Disease Classification
por: Sobuj, Md. Shohanur Islam, et al.
Publicado: (2024)
por: Sobuj, Md. Shohanur Islam, et al.
Publicado: (2024)
PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild
por: Yuan, Kun, et al.
Publicado: (2024)
por: Yuan, Kun, et al.
Publicado: (2024)
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
por: Tang, Longxiang, et al.
Publicado: (2024)
por: Tang, Longxiang, et al.
Publicado: (2024)
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
por: Guo, Wenliang, et al.
Publicado: (2025)
por: Guo, Wenliang, et al.
Publicado: (2025)
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
por: Chen, Jingxi, et al.
Publicado: (2024)
por: Chen, Jingxi, et al.
Publicado: (2024)
Generic Knowledge Boosted Pre-training For Remote Sensing Images
por: Huang, Ziyue, et al.
Publicado: (2024)
por: Huang, Ziyue, et al.
Publicado: (2024)
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
por: Dong, Xingning, et al.
Publicado: (2024)
por: Dong, Xingning, et al.
Publicado: (2024)
From Videos to Conversations: Egocentric Instructions for Task Assistance
por: Aggarwal, Lavisha, et al.
Publicado: (2026)
por: Aggarwal, Lavisha, et al.
Publicado: (2026)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
por: Wang, Hengkang, et al.
Publicado: (2025)
por: Wang, Hengkang, et al.
Publicado: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2025)
por: Kazakos, Evangelos, et al.
Publicado: (2025)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
por: Seminara, Luigi, et al.
Publicado: (2025)
por: Seminara, Luigi, et al.
Publicado: (2025)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
por: Wang, Weicheng, et al.
Publicado: (2026)
por: Wang, Weicheng, et al.
Publicado: (2026)
Ejemplares similares
-
Exploring Efficient Foundational Multi-modal Models for Video Summarization
por: Samel, Karan, et al.
Publicado: (2024) -
On the Efficacy of Text-Based Input Modalities for Action Anticipation
por: Beedu, Apoorva, et al.
Publicado: (2024) -
HierSum: A Global and Local Attention Mechanism for Video Summarization
por: Beedu, Apoorva, et al.
Publicado: (2025) -
Efficient Pre-training for Localized Instruction Generation of Videos
por: Batra, Anil, et al.
Publicado: (2023) -
SLAIM: Robust Dense Neural SLAM for Online Tracking and Mapping
por: Cartillier, Vincent, et al.
Publicado: (2024)