Predicting Implicit Arguments in Procedural Video Instructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Batra, Anil, Sevilla-Lara, Laura, Rohrbach, Marcus, Keller, Frank |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Pre-training for Localized Instruction Generation of Videos
von: Batra, Anil, et al.
Veröffentlicht: (2023)
von: Batra, Anil, et al.
Veröffentlicht: (2023)
Chrono: A Simple Blueprint for Representing Time in MLLMs
von: Rodriguez, Hector, et al.
Veröffentlicht: (2024)
von: Rodriguez, Hector, et al.
Veröffentlicht: (2024)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
von: Braun, Tobias, et al.
Veröffentlicht: (2024)
von: Braun, Tobias, et al.
Veröffentlicht: (2024)
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2025)
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
Coarse or Fine? Recognising Action End States without Labels
von: Moltisanti, Davide, et al.
Veröffentlicht: (2024)
von: Moltisanti, Davide, et al.
Veröffentlicht: (2024)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
von: Wu, Te-Lin, et al.
Veröffentlicht: (2021)
How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos
von: Loginova, Olga, et al.
Veröffentlicht: (2026)
von: Loginova, Olga, et al.
Veröffentlicht: (2026)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
von: Rodriguez, Hector G., et al.
Veröffentlicht: (2026)
von: Rodriguez, Hector G., et al.
Veröffentlicht: (2026)
SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
von: Niu, Yulei, et al.
Veröffentlicht: (2024)
von: Niu, Yulei, et al.
Veröffentlicht: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2026)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2026)
Open-Event Procedure Planning in Instructional Videos
von: Wu, Yilu, et al.
Veröffentlicht: (2024)
von: Wu, Yilu, et al.
Veröffentlicht: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
RECIPE: Procedural Planning via Grounding in Instructional Video
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
von: Wang, Hanlin, et al.
Veröffentlicht: (2023)
von: Wang, Hanlin, et al.
Veröffentlicht: (2023)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
von: Rothermel, Mark, et al.
Veröffentlicht: (2026)
von: Rothermel, Mark, et al.
Veröffentlicht: (2026)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
von: Zhou, Yufan, et al.
Veröffentlicht: (2025)
von: Zhou, Yufan, et al.
Veröffentlicht: (2025)
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
von: Arora, Aditya, et al.
Veröffentlicht: (2026)
von: Arora, Aditya, et al.
Veröffentlicht: (2026)
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
von: Hasegawa, Kimihiro, et al.
Veröffentlicht: (2025)
von: Hasegawa, Kimihiro, et al.
Veröffentlicht: (2025)
Bias for Action: Video Implicit Neural Representations with Bias Modulation
von: Kayabasi, Alper, et al.
Veröffentlicht: (2025)
von: Kayabasi, Alper, et al.
Veröffentlicht: (2025)
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
von: Toschi, Federico, et al.
Veröffentlicht: (2026)
von: Toschi, Federico, et al.
Veröffentlicht: (2026)
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
von: Huang, Xiyang, et al.
Veröffentlicht: (2026)
von: Huang, Xiyang, et al.
Veröffentlicht: (2026)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
VIMI: Grounding Video Generation through Multi-modal Instruction
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
Instruction Makes a Difference
von: Adewumi, Tosin, et al.
Veröffentlicht: (2024)
von: Adewumi, Tosin, et al.
Veröffentlicht: (2024)
Telling Stories for Common Sense Zero-Shot Action Recognition
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2023)
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2023)
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
von: Haneji, Yuto, et al.
Veröffentlicht: (2024)
von: Haneji, Yuto, et al.
Veröffentlicht: (2024)
Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
von: Li, Gen, et al.
Veröffentlicht: (2025)
von: Li, Gen, et al.
Veröffentlicht: (2025)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
von: Xing, Zhen, et al.
Veröffentlicht: (2024)
von: Xing, Zhen, et al.
Veröffentlicht: (2024)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models
von: Belouadi, Jonas, et al.
Veröffentlicht: (2025)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2025)
CI w/o TN: Context Injection without Task Name for Procedure Planning
von: Li, Xinjie
Veröffentlicht: (2024)
von: Li, Xinjie
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Pre-training for Localized Instruction Generation of Videos
von: Batra, Anil, et al.
Veröffentlicht: (2023) -
Chrono: A Simple Blueprint for Representing Time in MLLMs
von: Rodriguez, Hector, et al.
Veröffentlicht: (2024) -
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
von: Braun, Tobias, et al.
Veröffentlicht: (2024) -
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
von: Dagan, Gautier, et al.
Veröffentlicht: (2024) -
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2025)