ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Kaixin, Meng, Ziyang, Lin, Hongzhan, Luo, Ziyang, Tian, Yuchen, Ma, Jing, Huang, Zhiyong, Chua, Tat-Seng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-Tuning
di: Meng, Ziyang, et al.
Pubblicazione: (2024)
di: Meng, Ziyang, et al.
Pubblicazione: (2024)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
di: Tu, Songjun, et al.
Pubblicazione: (2025)
di: Tu, Songjun, et al.
Pubblicazione: (2025)
Does CLIP perceive art the same way we do?
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
di: Ouyang, Rongxin, et al.
Pubblicazione: (2024)
di: Ouyang, Rongxin, et al.
Pubblicazione: (2024)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
di: Zhang, Junbin, et al.
Pubblicazione: (2026)
di: Zhang, Junbin, et al.
Pubblicazione: (2026)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
di: Korolkov, Vasilii
Pubblicazione: (2025)
di: Korolkov, Vasilii
Pubblicazione: (2025)
Meaning over Motion: A Semantic-First Approach to 360° Viewport Prediction
di: Khah, Arman Nik, et al.
Pubblicazione: (2026)
di: Khah, Arman Nik, et al.
Pubblicazione: (2026)
Open High-Resolution Satellite Imagery: The WorldStrat Dataset -- With Application to Super-Resolution
di: Cornebise, Julien, et al.
Pubblicazione: (2022)
di: Cornebise, Julien, et al.
Pubblicazione: (2022)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
di: Cai, Weibin, et al.
Pubblicazione: (2025)
di: Cai, Weibin, et al.
Pubblicazione: (2025)
Efficient and Privacy-Protecting Background Removal for 2D Video Streaming using iPhone 15 Pro Max LiDAR
di: Kinnevan, Jessica, et al.
Pubblicazione: (2025)
di: Kinnevan, Jessica, et al.
Pubblicazione: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
di: Koh, Hyunseo, et al.
Pubblicazione: (2026)
di: Koh, Hyunseo, et al.
Pubblicazione: (2026)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
di: Qi, Dekang, et al.
Pubblicazione: (2026)
di: Qi, Dekang, et al.
Pubblicazione: (2026)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
ForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection
di: Samson, Hema Hariharan
Pubblicazione: (2026)
di: Samson, Hema Hariharan
Pubblicazione: (2026)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2024)
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2024)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
di: Mohammad, Noor Islam S.
Pubblicazione: (2025)
di: Mohammad, Noor Islam S.
Pubblicazione: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
di: Zhao, Lepeng, et al.
Pubblicazione: (2026)
di: Zhao, Lepeng, et al.
Pubblicazione: (2026)
Supervised Learning Has a Necessary Geometric Blind Spot: Theory, Consequences, and Minimal Repair
di: Rajput, Vishal
Pubblicazione: (2026)
di: Rajput, Vishal
Pubblicazione: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
di: Ma, Yueen, et al.
Pubblicazione: (2024)
di: Ma, Yueen, et al.
Pubblicazione: (2024)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
di: Alanazi, Ahmed, et al.
Pubblicazione: (2025)
di: Alanazi, Ahmed, et al.
Pubblicazione: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
di: Menon, Anjali R., et al.
Pubblicazione: (2025)
di: Menon, Anjali R., et al.
Pubblicazione: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
Step-Aware Residual-Guided Diffusion for EEG Spatial Super-Resolution
di: Liu, Hongjun, et al.
Pubblicazione: (2025)
di: Liu, Hongjun, et al.
Pubblicazione: (2025)
Banana Ripeness Level Classification using a Simple CNN Model Trained with Real and Synthetic Datasets
di: Chuquimarca, Luis, et al.
Pubblicazione: (2025)
di: Chuquimarca, Luis, et al.
Pubblicazione: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
di: Adžemović, Momir
Pubblicazione: (2025)
di: Adžemović, Momir
Pubblicazione: (2025)
Deterministic Event-Graph Substrates as World Models for Counterfactual Reasoning
di: Rovai, Fabio
Pubblicazione: (2026)
di: Rovai, Fabio
Pubblicazione: (2026)
DeepFusionNet: Autoencoder-Based Low-Light Image Enhancement and Super-Resolution
di: Çalışkan, Halil Hüseyin, et al.
Pubblicazione: (2025)
di: Çalışkan, Halil Hüseyin, et al.
Pubblicazione: (2025)
A Landmark-Aware Visual Navigation Dataset
di: Johnson, Faith, et al.
Pubblicazione: (2024)
di: Johnson, Faith, et al.
Pubblicazione: (2024)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
di: Sharma, Aditya, et al.
Pubblicazione: (2025)
di: Sharma, Aditya, et al.
Pubblicazione: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
di: Farzulla, Murad
Pubblicazione: (2026)
di: Farzulla, Murad
Pubblicazione: (2026)
Revisiting Energy-Based Model for Out-of-Distribution Detection
di: Wu, Yifan, et al.
Pubblicazione: (2024)
di: Wu, Yifan, et al.
Pubblicazione: (2024)
Progressive Cross Attention Network for Flood Segmentation using Multispectral Satellite Imagery
di: Feliren, Vicky, et al.
Pubblicazione: (2025)
di: Feliren, Vicky, et al.
Pubblicazione: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models
di: Yao, Dongyu, et al.
Pubblicazione: (2024)
di: Yao, Dongyu, et al.
Pubblicazione: (2024)
Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention
di: Korolkov, Vasilii, et al.
Pubblicazione: (2025)
di: Korolkov, Vasilii, et al.
Pubblicazione: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
di: Chen, Kewei, et al.
Pubblicazione: (2025)
di: Chen, Kewei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-Tuning
di: Meng, Ziyang, et al.
Pubblicazione: (2024) -
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
di: Zhang, Junbin, et al.
Pubblicazione: (2022) -
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
di: Tu, Songjun, et al.
Pubblicazione: (2025) -
Does CLIP perceive art the same way we do?
di: Asperti, Andrea, et al.
Pubblicazione: (2025) -
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
di: Ouyang, Rongxin, et al.
Pubblicazione: (2024)