ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seminara, Luigi, Moltisanti, Davide, Furnari, Antonino |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RECIPE: Procedural Planning via Grounding in Instructional Video
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
von: Seminara, Luigi, et al.
Veröffentlicht: (2024)
von: Seminara, Luigi, et al.
Veröffentlicht: (2024)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
von: Seminara, Luigi, et al.
Veröffentlicht: (2025)
von: Seminara, Luigi, et al.
Veröffentlicht: (2025)
Open-Event Procedure Planning in Instructional Videos
von: Wu, Yilu, et al.
Veröffentlicht: (2024)
von: Wu, Yilu, et al.
Veröffentlicht: (2024)
Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
von: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Veröffentlicht: (2024)
von: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Veröffentlicht: (2024)
ProSkill: Segment-Level Skill Assessment in Procedural Videos
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2026)
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2026)
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
von: Wang, Hanlin, et al.
Veröffentlicht: (2023)
von: Wang, Hanlin, et al.
Veröffentlicht: (2023)
LAP: A Language-Aware Planning Model For Procedure Planning In Instructional Videos
von: Shi, Lei, et al.
Veröffentlicht: (2026)
von: Shi, Lei, et al.
Veröffentlicht: (2026)
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
von: Zhou, Yufan, et al.
Veröffentlicht: (2025)
von: Zhou, Yufan, et al.
Veröffentlicht: (2025)
Calisthenics Skills Temporal Video Segmentation
von: Finocchiaro, Antonio, et al.
Veröffentlicht: (2025)
von: Finocchiaro, Antonio, et al.
Veröffentlicht: (2025)
Efficient Pre-training for Localized Instruction Generation of Videos
von: Batra, Anil, et al.
Veröffentlicht: (2023)
von: Batra, Anil, et al.
Veröffentlicht: (2023)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
von: Shi, Lei, et al.
Veröffentlicht: (2024)
von: Shi, Lei, et al.
Veröffentlicht: (2024)
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
von: Catinello, Alessandro Sebastiano, et al.
Veröffentlicht: (2025)
von: Catinello, Alessandro Sebastiano, et al.
Veröffentlicht: (2025)
Exploring Multimodal LMMs for Online Episodic Memory Question Answering on the Edge
von: Lando, Giuseppe, et al.
Veröffentlicht: (2026)
von: Lando, Giuseppe, et al.
Veröffentlicht: (2026)
EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision
von: Forte, Rosario, et al.
Veröffentlicht: (2026)
von: Forte, Rosario, et al.
Veröffentlicht: (2026)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2024)
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2024)
Continual Learning Improves Zero-Shot Action Recognition
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2024)
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2024)
RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
von: Zare, Ali, et al.
Veröffentlicht: (2024)
von: Zare, Ali, et al.
Veröffentlicht: (2024)
Efficient Calisthenics Skills Classification through Foreground Instance Selection and Depth Estimation
von: Finocchiaro, Antonio, et al.
Veröffentlicht: (2025)
von: Finocchiaro, Antonio, et al.
Veröffentlicht: (2025)
StillFast: An End-to-End Approach for Short-Term Object Interaction Anticipation
von: Ragusa, Francesco, et al.
Veröffentlicht: (2023)
von: Ragusa, Francesco, et al.
Veröffentlicht: (2023)
Coarse or Fine? Recognising Action End States without Labels
von: Moltisanti, Davide, et al.
Veröffentlicht: (2024)
von: Moltisanti, Davide, et al.
Veröffentlicht: (2024)
SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
von: Niu, Yulei, et al.
Veröffentlicht: (2024)
von: Niu, Yulei, et al.
Veröffentlicht: (2024)
Semantically Guided Action Anticipation
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
von: Samel, Karan, et al.
Veröffentlicht: (2025)
von: Samel, Karan, et al.
Veröffentlicht: (2025)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
von: Quattrocchi, Camillo, et al.
Veröffentlicht: (2023)
von: Quattrocchi, Camillo, et al.
Veröffentlicht: (2023)
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?
von: Lando, Giuseppe, et al.
Veröffentlicht: (2025)
von: Lando, Giuseppe, et al.
Veröffentlicht: (2025)
Exploiting Multimodal Synthetic Data for Egocentric Human-Object Interaction Detection in an Industrial Scenario
von: Leonardi, Rosario, et al.
Veröffentlicht: (2023)
von: Leonardi, Rosario, et al.
Veröffentlicht: (2023)
Are Synthetic Data Useful for Egocentric Hand-Object Interaction Detection?
von: Leonardi, Rosario, et al.
Veröffentlicht: (2023)
von: Leonardi, Rosario, et al.
Veröffentlicht: (2023)
FLASH Viterbi: Fast and Adaptive Viterbi Decoding for Modern Data Systems
von: Deng, Ziheng, et al.
Veröffentlicht: (2025)
von: Deng, Ziheng, et al.
Veröffentlicht: (2025)
How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction
von: Jun, Sejoon, et al.
Veröffentlicht: (2026)
von: Jun, Sejoon, et al.
Veröffentlicht: (2026)
Predicting Implicit Arguments in Procedural Video Instructions
von: Batra, Anil, et al.
Veröffentlicht: (2025)
von: Batra, Anil, et al.
Veröffentlicht: (2025)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
von: Rodin, Ivan, et al.
Veröffentlicht: (2025)
von: Rodin, Ivan, et al.
Veröffentlicht: (2025)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
Plan-X: Instruct Video Generation via Semantic Planning
von: Huang, Lun, et al.
Veröffentlicht: (2025)
von: Huang, Lun, et al.
Veröffentlicht: (2025)
RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
Quantum Viterbi Algorithm
von: Accardi, Luigi, et al.
Veröffentlicht: (2026)
von: Accardi, Luigi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RECIPE: Procedural Planning via Grounding in Instructional Video
von: Seminara, Luigi, et al.
Veröffentlicht: (2026) -
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
von: Seminara, Luigi, et al.
Veröffentlicht: (2024) -
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
von: Seminara, Luigi, et al.
Veröffentlicht: (2025) -
Open-Event Procedure Planning in Instructional Videos
von: Wu, Yilu, et al.
Veröffentlicht: (2024) -
Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
von: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Veröffentlicht: (2024)