From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Toschi, Federico, Brunello, Nicolò, Sassella, Andrea, Scotti, Vincenzo, Carman, Mark James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are complicated loss functions necessary for teaching LLMs to reason?
by: Carrino, Gabriele, et al.
Published: (2026)
by: Carrino, Gabriele, et al.
Published: (2026)
InTraVisTo: Inside Transformer Visualisation Tool
by: Brunello, Nicolò, et al.
Published: (2025)
by: Brunello, Nicolò, et al.
Published: (2025)
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
by: Singh, Raul, et al.
Published: (2025)
by: Singh, Raul, et al.
Published: (2025)
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
by: Sassella, Andrea, et al.
Published: (2026)
by: Sassella, Andrea, et al.
Published: (2026)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)
by: Wu, Te-Lin, et al.
Published: (2021)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
From Videos to Conversations: Egocentric Instructions for Task Assistance
by: Aggarwal, Lavisha, et al.
Published: (2026)
by: Aggarwal, Lavisha, et al.
Published: (2026)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
by: Wu, Chenyuan, et al.
Published: (2025)
by: Wu, Chenyuan, et al.
Published: (2025)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
by: Liang, Jiafeng, et al.
Published: (2024)
by: Liang, Jiafeng, et al.
Published: (2024)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
by: Nazir, Maham, et al.
Published: (2026)
by: Nazir, Maham, et al.
Published: (2026)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
by: Zhang, Kai, et al.
Published: (2023)
by: Zhang, Kai, et al.
Published: (2023)
Gazelle: An Instruction Dataset for Arabic Writing Assistance
by: Magdy, Samar M., et al.
Published: (2024)
by: Magdy, Samar M., et al.
Published: (2024)
Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
by: Aggarwal, Lavisha, et al.
Published: (2025)
by: Aggarwal, Lavisha, et al.
Published: (2025)
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
by: Horawalavithana, Sameera, et al.
Published: (2023)
by: Horawalavithana, Sameera, et al.
Published: (2023)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
Maya: An Instruction Finetuned Multilingual Multimodal Model
by: Alam, Nahid, et al.
Published: (2024)
by: Alam, Nahid, et al.
Published: (2024)
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
by: Hasegawa, Kimihiro, et al.
Published: (2025)
by: Hasegawa, Kimihiro, et al.
Published: (2025)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
by: Baraldi, Lorenzo, et al.
Published: (2025)
by: Baraldi, Lorenzo, et al.
Published: (2025)
Predicting Implicit Arguments in Procedural Video Instructions
by: Batra, Anil, et al.
Published: (2025)
by: Batra, Anil, et al.
Published: (2025)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
Instruction-tuning Aligns LLMs to the Human Brain
by: Aw, Khai Loong, et al.
Published: (2023)
by: Aw, Khai Loong, et al.
Published: (2023)
SemiHVision: Enhancing Medical Multimodal Models with a Semi-Human Annotated Dataset and Fine-Tuned Instruction Generation
by: Wang, Junda, et al.
Published: (2024)
by: Wang, Junda, et al.
Published: (2024)
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
by: Nguyen, Pha, et al.
Published: (2025)
by: Nguyen, Pha, et al.
Published: (2025)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
by: Du, Yifan, et al.
Published: (2023)
by: Du, Yifan, et al.
Published: (2023)
Instruction-Following Evaluation of Large Vision-Language Models
by: Shiono, Daiki, et al.
Published: (2025)
by: Shiono, Daiki, et al.
Published: (2025)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
by: Wang, Yanan, et al.
Published: (2025)
by: Wang, Yanan, et al.
Published: (2025)
VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
by: Glória-Silva, Diogo, et al.
Published: (2026)
by: Glória-Silva, Diogo, et al.
Published: (2026)
Leveraging LLMs for On-the-Fly Instruction Guided Image Editing
by: Santos, Rodrigo, et al.
Published: (2024)
by: Santos, Rodrigo, et al.
Published: (2024)
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
by: Biswas, Shristi Das, et al.
Published: (2026)
by: Biswas, Shristi Das, et al.
Published: (2026)
Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
by: Huang, Yupan, et al.
Published: (2023)
by: Huang, Yupan, et al.
Published: (2023)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
by: Tanaka, Ryota, et al.
Published: (2024)
by: Tanaka, Ryota, et al.
Published: (2024)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025)
by: Wu, Keming, et al.
Published: (2025)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
by: Cocchi, Federico, et al.
Published: (2025)
by: Cocchi, Federico, et al.
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
Similar Items
-
Are complicated loss functions necessary for teaching LLMs to reason?
by: Carrino, Gabriele, et al.
Published: (2026) -
InTraVisTo: Inside Transformer Visualisation Tool
by: Brunello, Nicolò, et al.
Published: (2025) -
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
by: Singh, Raul, et al.
Published: (2025) -
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
by: Sassella, Andrea, et al.
Published: (2026) -
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)