VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Glória-Silva, Diogo, Semedo, David, Maglhães, João |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Show and Guide: Instructional-Plan Grounded Vision and Language Model
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
Zero-Shot Action Recognition in Surveillance Videos
von: Pereira, Joao, et al.
Veröffentlicht: (2024)
von: Pereira, Joao, et al.
Veröffentlicht: (2024)
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
von: Pereira, João, et al.
Veröffentlicht: (2026)
von: Pereira, João, et al.
Veröffentlicht: (2026)
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
von: Pereira, Joao, et al.
Veröffentlicht: (2025)
von: Pereira, Joao, et al.
Veröffentlicht: (2025)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
von: Le, Huy, et al.
Veröffentlicht: (2025)
von: Le, Huy, et al.
Veröffentlicht: (2025)
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
von: Chaturvedi, Saket S., et al.
Veröffentlicht: (2025)
von: Chaturvedi, Saket S., et al.
Veröffentlicht: (2025)
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
von: Imrattanatrai, Wiradee, et al.
Veröffentlicht: (2025)
von: Imrattanatrai, Wiradee, et al.
Veröffentlicht: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
von: Li, Qian, et al.
Veröffentlicht: (2024)
von: Li, Qian, et al.
Veröffentlicht: (2024)
Predicting Implicit Arguments in Procedural Video Instructions
von: Batra, Anil, et al.
Veröffentlicht: (2025)
von: Batra, Anil, et al.
Veröffentlicht: (2025)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
Leveraging LLMs for On-the-Fly Instruction Guided Image Editing
von: Santos, Rodrigo, et al.
Veröffentlicht: (2024)
von: Santos, Rodrigo, et al.
Veröffentlicht: (2024)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
von: Kriz, Reno, et al.
Veröffentlicht: (2024)
von: Kriz, Reno, et al.
Veröffentlicht: (2024)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2023)
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2023)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
AutoArabic: A Three-Stage Framework for Localizing Video-Text Retrieval Benchmarks
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance
von: Venkatesh, Kavana, et al.
Veröffentlicht: (2024)
von: Venkatesh, Kavana, et al.
Veröffentlicht: (2024)
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
von: Fu, Zihang, et al.
Veröffentlicht: (2026)
von: Fu, Zihang, et al.
Veröffentlicht: (2026)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
Fostering Video Reasoning via Next-Event Prediction
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
Beyond Coarse-Grained Matching in Video-Text Retrieval
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
von: Chen, Aozhu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Show and Guide: Instructional-Plan Grounded Vision and Language Model
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024) -
Zero-Shot Action Recognition in Surveillance Videos
von: Pereira, Joao, et al.
Veröffentlicht: (2024) -
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
von: Pereira, João, et al.
Veröffentlicht: (2026) -
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
von: Pereira, Joao, et al.
Veröffentlicht: (2025) -
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
von: Le, Huy, et al.
Veröffentlicht: (2025)