Computer User Interface Understanding. A New Dataset and a Learning Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Muñoz, Andrés, Borrajo, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Veri-Car: Towards Open-world Vehicle Information Retrieval
von: Muñoz, Andrés, et al.
Veröffentlicht: (2024)
von: Muñoz, Andrés, et al.
Veröffentlicht: (2024)
Hypercone Assisted Contour Generation for Out-of-Distribution Detection
von: Vapsi, Annita, et al.
Veröffentlicht: (2025)
von: Vapsi, Annita, et al.
Veröffentlicht: (2025)
DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2024)
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2024)
A Low-Computational Video Synopsis Framework with a Standard Dataset
von: Malekpour, Ramtin, et al.
Veröffentlicht: (2024)
von: Malekpour, Ramtin, et al.
Veröffentlicht: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
Towards Effective Image Forensics via A Novel Computationally Efficient Framework and A New Image Splice Dataset
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
A Lightweight Vision-Language Fusion Framework for Predicting App Ratings from User Interfaces and Metadata
von: Sultana, Azrin, et al.
Veröffentlicht: (2026)
von: Sultana, Azrin, et al.
Veröffentlicht: (2026)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
A Deep Learning Framework for Visual Attention Prediction and Analysis of News Interfaces
von: Kenely, Matthew, et al.
Veröffentlicht: (2025)
von: Kenely, Matthew, et al.
Veröffentlicht: (2025)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
A New Dataset and Framework for Real-World Blurred Images Super-Resolution
von: Qin, Rui, et al.
Veröffentlicht: (2024)
von: Qin, Rui, et al.
Veröffentlicht: (2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
von: Li, Zhangheng, et al.
Veröffentlicht: (2024)
Collection Space Navigator: An Interactive Visualization Interface for Multidimensional Datasets
von: Ohm, Tillmann, et al.
Veröffentlicht: (2023)
von: Ohm, Tillmann, et al.
Veröffentlicht: (2023)
Interactive Interface For Semantic Segmentation Dataset Synthesis
von: Tran, Ngoc-Do, et al.
Veröffentlicht: (2025)
von: Tran, Ngoc-Do, et al.
Veröffentlicht: (2025)
ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
von: Gao, Difei, et al.
Veröffentlicht: (2023)
von: Gao, Difei, et al.
Veröffentlicht: (2023)
Shining a Light on Hurricane Damage Estimation via Nighttime Light Data: Pre-processing Matters
von: Thomas, Nancy, et al.
Veröffentlicht: (2024)
von: Thomas, Nancy, et al.
Veröffentlicht: (2024)
CBEN -- A Multimodal Machine Learning Dataset for Cloud Robust Remote Sensing Image Understanding
von: Stricker, Marco, et al.
Veröffentlicht: (2026)
von: Stricker, Marco, et al.
Veröffentlicht: (2026)
VUDG: A Dataset for Video Understanding Domain Generalization
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding
von: Yuan, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Yuan, Xiaoyu, et al.
Veröffentlicht: (2026)
FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos
von: Wu, Zhengqian, et al.
Veröffentlicht: (2024)
von: Wu, Zhengqian, et al.
Veröffentlicht: (2024)
SkyScenes: A Synthetic Dataset for Aerial Scene Understanding
von: Khose, Sahil, et al.
Veröffentlicht: (2023)
von: Khose, Sahil, et al.
Veröffentlicht: (2023)
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset
von: Waghmare, Sagar M., et al.
Veröffentlicht: (2023)
von: Waghmare, Sagar M., et al.
Veröffentlicht: (2023)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
von: Chen, Guo, et al.
Veröffentlicht: (2026)
von: Chen, Guo, et al.
Veröffentlicht: (2026)
Surgical Visual Understanding (SurgVU) Dataset
von: Zia, Aneeq, et al.
Veröffentlicht: (2025)
von: Zia, Aneeq, et al.
Veröffentlicht: (2025)
Exploring the Effect of Dataset Diversity in Self-Supervised Learning for Surgical Computer Vision
von: Jaspers, Tim J. M., et al.
Veröffentlicht: (2024)
von: Jaspers, Tim J. M., et al.
Veröffentlicht: (2024)
Infrared Small Target Detection in Satellite Videos: A New Dataset and A Novel Recurrent Feature Refinement Framework
von: Ying, Xinyi, et al.
Veröffentlicht: (2024)
von: Ying, Xinyi, et al.
Veröffentlicht: (2024)
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework
von: Hu, Yutong, et al.
Veröffentlicht: (2026)
von: Hu, Yutong, et al.
Veröffentlicht: (2026)
Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning
von: Zhu, Ye, et al.
Veröffentlicht: (2025)
von: Zhu, Ye, et al.
Veröffentlicht: (2025)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
von: Xie, Tianbao, et al.
Veröffentlicht: (2025)
von: Xie, Tianbao, et al.
Veröffentlicht: (2025)
CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models
von: Pyatov, Vladislav, et al.
Veröffentlicht: (2026)
von: Pyatov, Vladislav, et al.
Veröffentlicht: (2026)
RSUD20K: A Dataset for Road Scene Understanding In Autonomous Driving
von: Zunair, Hasib, et al.
Veröffentlicht: (2024)
von: Zunair, Hasib, et al.
Veröffentlicht: (2024)
GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
von: Ferrod, Roger, et al.
Veröffentlicht: (2026)
von: Ferrod, Roger, et al.
Veröffentlicht: (2026)
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
von: Chua, Phoebe, et al.
Veröffentlicht: (2025)
von: Chua, Phoebe, et al.
Veröffentlicht: (2025)
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
von: Peddi, Rohith, et al.
Veröffentlicht: (2023)
von: Peddi, Rohith, et al.
Veröffentlicht: (2023)
Enabling Advanced Land Cover Analytics: An Integrated Data Extraction Pipeline for Predictive Modeling with the Dynamic World Dataset
von: Radermecker, Victor, et al.
Veröffentlicht: (2024)
von: Radermecker, Victor, et al.
Veröffentlicht: (2024)
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
AutoFormBench: Benchmark Dataset for Automating Form Understanding
von: Baral, Gaurab, et al.
Veröffentlicht: (2026)
von: Baral, Gaurab, et al.
Veröffentlicht: (2026)
Advancing Video Anomaly Detection: A Concise Review and a New Dataset
von: Zhu, Liyun, et al.
Veröffentlicht: (2024)
von: Zhu, Liyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Veri-Car: Towards Open-world Vehicle Information Retrieval
von: Muñoz, Andrés, et al.
Veröffentlicht: (2024) -
Hypercone Assisted Contour Generation for Out-of-Distribution Detection
von: Vapsi, Annita, et al.
Veröffentlicht: (2025) -
DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2024) -
A Low-Computational Video Synopsis Framework with a Standard Dataset
von: Malekpour, Ramtin, et al.
Veröffentlicht: (2024) -
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025)