Generative Timelines for Instructed Visual Assembly

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pardo, Alejandro, Wang, Jui-Hsien, Ghanem, Bernard, Sivic, Josef, Russell, Bryan, Heilbron, Fabian Caba
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929597729210368
author Pardo, Alejandro
Wang, Jui-Hsien
Ghanem, Bernard
Sivic, Josef
Russell, Bryan
Heilbron, Fabian Caba
author_facet Pardo, Alejandro
Wang, Jui-Hsien
Ghanem, Bernard
Sivic, Josef
Russell, Bryan
Heilbron, Fabian Caba
contents The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task Instructed visual assembly. This task is challenging as it requires (i) identifying relevant visual content in the input timeline as well as retrieving relevant visual content in a given input (video) collection, (ii) understanding the input natural language instruction, and (iii) performing the desired edits of the input visual timeline to produce an output timeline. To address these challenges, we propose the Timeline Assembler, a generative model trained to perform instructed visual assembly tasks. The contributions of this work are three-fold. First, we develop a large multimodal language model, which is designed to process visual content, compactly represent timelines and accurately interpret timeline editing instructions. Second, we introduce a novel method for automatically generating datasets for visual assembly tasks, enabling efficient training of our model without the need for human-labeled data. Third, we validate our approach by creating two novel datasets for image and video assembly, demonstrating that the Timeline Assembler substantially outperforms established baseline models, including the recent GPT-4o, in accurately executing complex assembly instructions across various real-world inspired scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12293
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generative Timelines for Instructed Visual Assembly
Pardo, Alejandro
Wang, Jui-Hsien
Ghanem, Bernard
Sivic, Josef
Russell, Bryan
Heilbron, Fabian Caba
Computer Vision and Pattern Recognition
Human-Computer Interaction
Multimedia
The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task Instructed visual assembly. This task is challenging as it requires (i) identifying relevant visual content in the input timeline as well as retrieving relevant visual content in a given input (video) collection, (ii) understanding the input natural language instruction, and (iii) performing the desired edits of the input visual timeline to produce an output timeline. To address these challenges, we propose the Timeline Assembler, a generative model trained to perform instructed visual assembly tasks. The contributions of this work are three-fold. First, we develop a large multimodal language model, which is designed to process visual content, compactly represent timelines and accurately interpret timeline editing instructions. Second, we introduce a novel method for automatically generating datasets for visual assembly tasks, enabling efficient training of our model without the need for human-labeled data. Third, we validate our approach by creating two novel datasets for image and video assembly, demonstrating that the Timeline Assembler substantially outperforms established baseline models, including the recent GPT-4o, in accurately executing complex assembly instructions across various real-world inspired scenarios.
title Generative Timelines for Instructed Visual Assembly
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
Multimedia
url https://arxiv.org/abs/2411.12293