Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jang, Yunseok, Song, Yeda, Sohn, Sungryull, Logeswaran, Lajanugen, Luo, Tiange, Kim, Dong-Ki, Bae, Kyunghoon, Lee, Honglak
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2505.12632
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918024363114496
author Jang, Yunseok
Song, Yeda
Sohn, Sungryull
Logeswaran, Lajanugen
Luo, Tiange
Kim, Dong-Ki
Bae, Kyunghoon
Lee, Honglak
author_facet Jang, Yunseok
Song, Yeda
Sohn, Sungryull
Logeswaran, Lajanugen
Luo, Tiange
Kim, Dong-Ki
Bae, Kyunghoon
Lee, Honglak
contents Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructional videos capturing diverse real-world mobile OS navigation across multiple platforms. Models that include MONDAY in their pre-training phases demonstrate robust cross-platform generalization capabilities, consistently outperforming models trained on existing single OS datasets while achieving an average performance gain of 18.11%p on an unseen mobile OS platform. To enable continuous dataset expansion as mobile platforms evolve, we present an automated framework that leverages publicly available video content to create comprehensive task datasets without manual annotation. Our framework comprises robust OCR-based scene detection (95.04% F1score), near-perfect UI element detection (99.87% hit ratio), and novel multi-step action identification to extract reliable action sequences across diverse interface configurations. We contribute both the MONDAY dataset and our automated collection framework to facilitate future research in mobile OS navigation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
Jang, Yunseok
Song, Yeda
Sohn, Sungryull
Logeswaran, Lajanugen
Luo, Tiange
Kim, Dong-Ki
Bae, Kyunghoon
Lee, Honglak
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructional videos capturing diverse real-world mobile OS navigation across multiple platforms. Models that include MONDAY in their pre-training phases demonstrate robust cross-platform generalization capabilities, consistently outperforming models trained on existing single OS datasets while achieving an average performance gain of 18.11%p on an unseen mobile OS platform. To enable continuous dataset expansion as mobile platforms evolve, we present an automated framework that leverages publicly available video content to create comprehensive task datasets without manual annotation. Our framework comprises robust OCR-based scene detection (95.04% F1score), near-perfect UI element detection (99.87% hit ratio), and novel multi-step action identification to extract reliable action sequences across diverse interface configurations. We contribute both the MONDAY dataset and our automated collection framework to facilitate future research in mobile OS navigation.
title Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.12632