V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rahman, Abdur, Chawla, Rajat, Kumar, Muskaan, Datta, Arkajit, Jha, Adarsh, NS, Mukunda, Bhola, Ishaan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913438963335168
author Rahman, Abdur
Chawla, Rajat
Kumar, Muskaan
Datta, Arkajit
Jha, Adarsh
NS, Mukunda
Bhola, Ishaan
author_facet Rahman, Abdur
Chawla, Rajat
Kumar, Muskaan
Datta, Arkajit
Jha, Adarsh
NS, Mukunda
Bhola, Ishaan
contents In the rapidly evolving landscape of AI research and application, Multimodal Large Language Models (MLLMs) have emerged as a transformative force, adept at interpreting and integrating information from diverse modalities such as text, images, and Graphical User Interfaces (GUIs). Despite these advancements, the nuanced interaction and understanding of GUIs pose a significant challenge, limiting the potential of existing models to enhance automation levels. To bridge this gap, this paper presents V-Zen, an innovative Multimodal Large Language Model (MLLM) meticulously crafted to revolutionise the domain of GUI understanding and grounding. Equipped with dual-resolution image encoders, V-Zen establishes new benchmarks in efficient grounding and next-action prediction, thereby laying the groundwork for self-operating computer systems. Complementing V-Zen is the GUIDE dataset, an extensive collection of real-world GUI elements and task-based sequences, serving as a catalyst for specialised fine-tuning. The successful integration of V-Zen and GUIDE marks the dawn of a new era in multimodal AI research, opening the door to intelligent, autonomous computing experiences. This paper extends an invitation to the research community to join this exciting journey, shaping the future of GUI automation. In the spirit of open science, our code, data, and model will be made publicly available, paving the way for multimodal dialogue scenarios with intricate and precise interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15341
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
Rahman, Abdur
Chawla, Rajat
Kumar, Muskaan
Datta, Arkajit
Jha, Adarsh
NS, Mukunda
Bhola, Ishaan
Artificial Intelligence
Computer Vision and Pattern Recognition
In the rapidly evolving landscape of AI research and application, Multimodal Large Language Models (MLLMs) have emerged as a transformative force, adept at interpreting and integrating information from diverse modalities such as text, images, and Graphical User Interfaces (GUIs). Despite these advancements, the nuanced interaction and understanding of GUIs pose a significant challenge, limiting the potential of existing models to enhance automation levels. To bridge this gap, this paper presents V-Zen, an innovative Multimodal Large Language Model (MLLM) meticulously crafted to revolutionise the domain of GUI understanding and grounding. Equipped with dual-resolution image encoders, V-Zen establishes new benchmarks in efficient grounding and next-action prediction, thereby laying the groundwork for self-operating computer systems. Complementing V-Zen is the GUIDE dataset, an extensive collection of real-world GUI elements and task-based sequences, serving as a catalyst for specialised fine-tuning. The successful integration of V-Zen and GUIDE marks the dawn of a new era in multimodal AI research, opening the door to intelligent, autonomous computing experiences. This paper extends an invitation to the research community to join this exciting journey, shaping the future of GUI automation. In the spirit of open science, our code, data, and model will be made publicly available, paving the way for multimodal dialogue scenarios with intricate and precise interactions.
title V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15341