Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Nagasinghe, Kumaranage Ravindu Yasas, Zhou, Honglu, Gunawardhana, Malitha, Min, Martin Renqiang, Harari, Daniel, Khan, Muhammad Haris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Effective are Self-Supervised Models for Contact Identification in Videos
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024)
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
di: Ahamed, Shihab Aaqil, et al.
Pubblicazione: (2025)
di: Ahamed, Shihab Aaqil, et al.
Pubblicazione: (2025)
Towards Generalizing to Unseen Domains with Few Labels
di: Galappaththige, Chamuditha Jayanga, et al.
Pubblicazione: (2024)
di: Galappaththige, Chamuditha Jayanga, et al.
Pubblicazione: (2024)
How good nnU-Net for Segmenting Cardiac MRI: A Comprehensive Evaluation
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024)
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024)
Integrating Deep Learning in Cardiology: A Comprehensive Review of Atrial Fibrillation, Left Atrial Scar Segmentation, and the Frontiers of State-of-the-Art Techniques
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024)
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024)
Segmenting Bi-Atrial Structures Using ResNext Based Framework
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2025)
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2025)
Knowledge-guided Continual Learning for Behavioral Analytics Systems
di: Senarath, Yasas, et al.
Pubblicazione: (2025)
di: Senarath, Yasas, et al.
Pubblicazione: (2025)
Domain-Guided Weight Modulation for Semi-Supervised Domain Generalization
di: Galappaththige, Chamuditha Jayanaga, et al.
Pubblicazione: (2024)
di: Galappaththige, Chamuditha Jayanaga, et al.
Pubblicazione: (2024)
Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
di: Harari, Daniel, et al.
Pubblicazione: (2025)
di: Harari, Daniel, et al.
Pubblicazione: (2025)
ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
di: Seminara, Luigi, et al.
Pubblicazione: (2026)
di: Seminara, Luigi, et al.
Pubblicazione: (2026)
Open-Event Procedure Planning in Instructional Videos
di: Wu, Yilu, et al.
Pubblicazione: (2024)
di: Wu, Yilu, et al.
Pubblicazione: (2024)
Performance of Recent Large Language Models for a Low-Resourced Language
di: Jayakody, Ravindu, et al.
Pubblicazione: (2024)
di: Jayakody, Ravindu, et al.
Pubblicazione: (2024)
LAP: A Language-Aware Planning Model For Procedure Planning In Instructional Videos
di: Shi, Lei, et al.
Pubblicazione: (2026)
di: Shi, Lei, et al.
Pubblicazione: (2026)
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
di: Zhou, Yufan, et al.
Pubblicazione: (2025)
di: Zhou, Yufan, et al.
Pubblicazione: (2025)
RECIPE: Procedural Planning via Grounding in Instructional Video
di: Seminara, Luigi, et al.
Pubblicazione: (2026)
di: Seminara, Luigi, et al.
Pubblicazione: (2026)
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
di: Wang, Hanlin, et al.
Pubblicazione: (2023)
di: Wang, Hanlin, et al.
Pubblicazione: (2023)
Dynamic Position Transformation and Boundary Refinement Network for Left Atrial Segmentation
di: Xu, Fangqiang, et al.
Pubblicazione: (2024)
di: Xu, Fangqiang, et al.
Pubblicazione: (2024)
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
di: Ding, Bonan, et al.
Pubblicazione: (2026)
di: Ding, Bonan, et al.
Pubblicazione: (2026)
Divergent Domains, Convergent Grading: Enhancing Generalization in Diabetic Retinopathy Grading
di: Chokuwa, Sharon, et al.
Pubblicazione: (2024)
di: Chokuwa, Sharon, et al.
Pubblicazione: (2024)
Key-Conditioned Orthonormal Transform Gating (K-OTG): Multi-Key Access Control with Hidden-State Scrambling for LoRA-Tuned Models
di: Khan, Muhammad Haris
Pubblicazione: (2025)
di: Khan, Muhammad Haris
Pubblicazione: (2025)
Is Monotonic Sampling Necessary in Diffusion Models?
di: Khan, Muhammad Haris
Pubblicazione: (2026)
di: Khan, Muhammad Haris
Pubblicazione: (2026)
SafeBench-Seq: A Homology-Clustered, CPU-Only Baseline for Protein Hazard Screening with Physicochemical/Composition Features and Cluster-Aware Confidence Intervals
di: Khan, Muhammad Haris
Pubblicazione: (2025)
di: Khan, Muhammad Haris
Pubblicazione: (2025)
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
di: Chen, Yuxiao, et al.
Pubblicazione: (2024)
di: Chen, Yuxiao, et al.
Pubblicazione: (2024)
RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
di: Zare, Ali, et al.
Pubblicazione: (2024)
di: Zare, Ali, et al.
Pubblicazione: (2024)
SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
di: Niu, Yulei, et al.
Pubblicazione: (2024)
di: Niu, Yulei, et al.
Pubblicazione: (2024)
Improving Pseudo-labelling and Enhancing Robustness for Semi-Supervised Domain Generalization
di: Khan, Adnan, et al.
Pubblicazione: (2024)
di: Khan, Adnan, et al.
Pubblicazione: (2024)
Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
di: Kuckreja, Kartik, et al.
Pubblicazione: (2026)
di: Kuckreja, Kartik, et al.
Pubblicazione: (2026)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
di: Shi, Lei, et al.
Pubblicazione: (2024)
di: Shi, Lei, et al.
Pubblicazione: (2024)
Robust and Label-Efficient Deep Waste Detection
di: Abid, Hassan, et al.
Pubblicazione: (2025)
di: Abid, Hassan, et al.
Pubblicazione: (2025)
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
di: Samel, Karan, et al.
Pubblicazione: (2025)
di: Samel, Karan, et al.
Pubblicazione: (2025)
CountZES: Counting via Zero-Shot Exemplar Selection
di: Siddiqui, Muhammad Ibraheem, et al.
Pubblicazione: (2025)
di: Siddiqui, Muhammad Ibraheem, et al.
Pubblicazione: (2025)
Noise-Tolerant Few-Shot Unsupervised Adapter for Vision-Language Models
di: Ali, Eman, et al.
Pubblicazione: (2023)
di: Ali, Eman, et al.
Pubblicazione: (2023)
Sharpen Your Skills: Textbook Format Braille.
di: Tate, Barbara H., et al.
Pubblicazione: (1982)
di: Tate, Barbara H., et al.
Pubblicazione: (1982)
Predicting Implicit Arguments in Procedural Video Instructions
di: Batra, Anil, et al.
Pubblicazione: (2025)
di: Batra, Anil, et al.
Pubblicazione: (2025)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
di: Maaz, Muhammad, et al.
Pubblicazione: (2024)
di: Maaz, Muhammad, et al.
Pubblicazione: (2024)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
di: Wong, Zhen Hao, et al.
Pubblicazione: (2025)
di: Wong, Zhen Hao, et al.
Pubblicazione: (2025)
VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping
di: Baliah, Sanoojan, et al.
Pubblicazione: (2026)
di: Baliah, Sanoojan, et al.
Pubblicazione: (2026)
Guidelines for the Review and Selection of Textbooks and Instructional Materials.
Pubblicazione: (1976)
Pubblicazione: (1976)
Downscaling model based on CorrDiff
di: Sun, Honglu
Pubblicazione: (2026)
di: Sun, Honglu
Pubblicazione: (2026)
Documenti analoghi
-
How Effective are Self-Supervised Models for Contact Identification in Videos
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024) -
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
di: Ahamed, Shihab Aaqil, et al.
Pubblicazione: (2025) -
Towards Generalizing to Unseen Domains with Few Labels
di: Galappaththige, Chamuditha Jayanga, et al.
Pubblicazione: (2024) -
How good nnU-Net for Segmenting Cardiac MRI: A Comprehensive Evaluation
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024) -
Integrating Deep Learning in Cardiology: A Comprehensive Review of Atrial Fibrillation, Left Atrial Scar Segmentation, and the Frontiers of State-of-the-Art Techniques
di: Gunawardhana, Malitha, et al.
Pubblicazione: (2024)