Saved in:
| Main Authors: | Muñoz, Andrés, Borrajo, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.10170 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Veri-Car: Towards Open-world Vehicle Information Retrieval
by: Muñoz, Andrés, et al.
Published: (2024)
by: Muñoz, Andrés, et al.
Published: (2024)
Hypercone Assisted Contour Generation for Out-of-Distribution Detection
by: Vapsi, Annita, et al.
Published: (2025)
by: Vapsi, Annita, et al.
Published: (2025)
DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
by: Peng, Yi-Hao, et al.
Published: (2024)
by: Peng, Yi-Hao, et al.
Published: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
A Low-Computational Video Synopsis Framework with a Standard Dataset
by: Malekpour, Ramtin, et al.
Published: (2024)
by: Malekpour, Ramtin, et al.
Published: (2024)
Computer-Use Agents as Judges for Generative User Interface
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
Towards Effective Image Forensics via A Novel Computationally Efficient Framework and A New Image Splice Dataset
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
A Deep Learning Framework for Visual Attention Prediction and Analysis of News Interfaces
by: Kenely, Matthew, et al.
Published: (2025)
by: Kenely, Matthew, et al.
Published: (2025)
Shining a Light on Hurricane Damage Estimation via Nighttime Light Data: Pre-processing Matters
by: Thomas, Nancy, et al.
Published: (2024)
by: Thomas, Nancy, et al.
Published: (2024)
Collection Space Navigator: An Interactive Visualization Interface for Multidimensional Datasets
by: Ohm, Tillmann, et al.
Published: (2023)
by: Ohm, Tillmann, et al.
Published: (2023)
A Lightweight Vision-Language Fusion Framework for Predicting App Ratings from User Interfaces and Metadata
by: Sultana, Azrin, et al.
Published: (2026)
by: Sultana, Azrin, et al.
Published: (2026)
Enabling Advanced Land Cover Analytics: An Integrated Data Extraction Pipeline for Predictive Modeling with the Dynamic World Dataset
by: Radermecker, Victor, et al.
Published: (2024)
by: Radermecker, Victor, et al.
Published: (2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
by: Li, Zhangheng, et al.
Published: (2024)
by: Li, Zhangheng, et al.
Published: (2024)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
A New Dataset and Framework for Real-World Blurred Images Super-Resolution
by: Qin, Rui, et al.
Published: (2024)
by: Qin, Rui, et al.
Published: (2024)
Interactive Interface For Semantic Segmentation Dataset Synthesis
by: Tran, Ngoc-Do, et al.
Published: (2025)
by: Tran, Ngoc-Do, et al.
Published: (2025)
ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
by: Gao, Difei, et al.
Published: (2023)
by: Gao, Difei, et al.
Published: (2023)
VUDG: A Dataset for Video Understanding Domain Generalization
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
by: Xie, Tianbao, et al.
Published: (2025)
by: Xie, Tianbao, et al.
Published: (2025)
CBEN -- A Multimodal Machine Learning Dataset for Cloud Robust Remote Sensing Image Understanding
by: Stricker, Marco, et al.
Published: (2026)
by: Stricker, Marco, et al.
Published: (2026)
MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding
by: Yuan, Xiaoyu, et al.
Published: (2026)
by: Yuan, Xiaoyu, et al.
Published: (2026)
FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos
by: Wu, Zhengqian, et al.
Published: (2024)
by: Wu, Zhengqian, et al.
Published: (2024)
SkyScenes: A Synthetic Dataset for Aerial Scene Understanding
by: Khose, Sahil, et al.
Published: (2023)
by: Khose, Sahil, et al.
Published: (2023)
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset
by: Waghmare, Sagar M., et al.
Published: (2023)
by: Waghmare, Sagar M., et al.
Published: (2023)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
by: Chen, Guo, et al.
Published: (2026)
by: Chen, Guo, et al.
Published: (2026)
Surgical Visual Understanding (SurgVU) Dataset
by: Zia, Aneeq, et al.
Published: (2025)
by: Zia, Aneeq, et al.
Published: (2025)
Exploring the Effect of Dataset Diversity in Self-Supervised Learning for Surgical Computer Vision
by: Jaspers, Tim J. M., et al.
Published: (2024)
by: Jaspers, Tim J. M., et al.
Published: (2024)
Infrared Small Target Detection in Satellite Videos: A New Dataset and A Novel Recurrent Feature Refinement Framework
by: Ying, Xinyi, et al.
Published: (2024)
by: Ying, Xinyi, et al.
Published: (2024)
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
A Chinese Multi-label Affective Computing Dataset Based on Social Media Network Users
by: Zhou, Jingyi, et al.
Published: (2024)
by: Zhou, Jingyi, et al.
Published: (2024)
CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models
by: Pyatov, Vladislav, et al.
Published: (2026)
by: Pyatov, Vladislav, et al.
Published: (2026)
Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework
by: Hu, Yutong, et al.
Published: (2026)
by: Hu, Yutong, et al.
Published: (2026)
Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning
by: Zhu, Ye, et al.
Published: (2025)
by: Zhu, Ye, et al.
Published: (2025)
A New Dataset and Framework for Robust Road Surface Classification via Camera-IMU Fusion
by: Costa, Willams de Lima, et al.
Published: (2026)
by: Costa, Willams de Lima, et al.
Published: (2026)
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
RSUD20K: A Dataset for Road Scene Understanding In Autonomous Driving
by: Zunair, Hasib, et al.
Published: (2024)
by: Zunair, Hasib, et al.
Published: (2024)
GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
by: Ferrod, Roger, et al.
Published: (2026)
by: Ferrod, Roger, et al.
Published: (2026)
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
by: Chua, Phoebe, et al.
Published: (2025)
by: Chua, Phoebe, et al.
Published: (2025)
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
by: Peddi, Rohith, et al.
Published: (2023)
by: Peddi, Rohith, et al.
Published: (2023)
Similar Items
-
Veri-Car: Towards Open-world Vehicle Information Retrieval
by: Muñoz, Andrés, et al.
Published: (2024) -
Hypercone Assisted Contour Generation for Out-of-Distribution Detection
by: Vapsi, Annita, et al.
Published: (2025) -
DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
by: Peng, Yi-Hao, et al.
Published: (2024) -
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025) -
A Low-Computational Video Synopsis Framework with a Standard Dataset
by: Malekpour, Ramtin, et al.
Published: (2024)