Aria-UI: Visual Grounding for GUI Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yuhao, Wang, Yue, Li, Dongxu, Luo, Ziyang, Chen, Bei, Huang, Chao, Li, Junnan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025)
by: Jing, Hongyi, et al.
Published: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025)
by: Qin, Yujia, et al.
Published: (2025)
MAIC-UI: Making Interactive Courseware with Generative UI
by: Tu, Shangqing, et al.
Published: (2026)
by: Tu, Shangqing, et al.
Published: (2026)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
GhostUI: Unveiling Hidden Interactions in Mobile UI
by: Kweon, Minkyu, et al.
Published: (2026)
by: Kweon, Minkyu, et al.
Published: (2026)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)
by: Tang, Liujian, et al.
Published: (2025)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
SPRITE: From Static Mockups to Engine-Ready Game UI
by: Bai, Yunshu, et al.
Published: (2026)
by: Bai, Yunshu, et al.
Published: (2026)
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
by: Zhang, Li, et al.
Published: (2024)
by: Zhang, Li, et al.
Published: (2024)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025)
by: Hu, Siyuan, et al.
Published: (2025)
HelpViz: Automatic Generation of Contextual Visual MobileTutorials from Text-Based Instructions
by: Zhong, Mingyuan, et al.
Published: (2021)
by: Zhong, Mingyuan, et al.
Published: (2021)
Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
by: Liao, Zeyi, et al.
Published: (2025)
by: Liao, Zeyi, et al.
Published: (2025)
AutoGameUI: Constructing High-Fidelity GameUI via Multimodal Correspondence Matching
by: Tang, Zhongliang, et al.
Published: (2024)
by: Tang, Zhongliang, et al.
Published: (2024)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
by: Sun, Jiazheng, et al.
Published: (2026)
by: Sun, Jiazheng, et al.
Published: (2026)
GUI Agents: A Survey
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
by: Liu, Zhilin, et al.
Published: (2026)
by: Liu, Zhilin, et al.
Published: (2026)
ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow
by: Garg, Aryan, et al.
Published: (2025)
by: Garg, Aryan, et al.
Published: (2025)
FEAD: Figma-Enhanced App Design Framework for Improving UI/UX in Educational App Development
by: Huang, Tianyi
Published: (2024)
by: Huang, Tianyi
Published: (2024)
Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding
by: Tan, Wee Joe, et al.
Published: (2026)
by: Tan, Wee Joe, et al.
Published: (2026)
UI-UG: A Unified MLLM for UI Understanding and Generation
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Magentic-UI: Towards Human-in-the-loop Agentic Systems
by: Mozannar, Hussein, et al.
Published: (2025)
by: Mozannar, Hussein, et al.
Published: (2025)
GUI Agents with Foundation Models: A Comprehensive Survey
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
by: Tang, Jingyu, et al.
Published: (2025)
by: Tang, Jingyu, et al.
Published: (2025)
Bridging Gulfs in UI Generation through Semantic Guidance
by: Park, Seokhyeon, et al.
Published: (2026)
by: Park, Seokhyeon, et al.
Published: (2026)
InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMs
by: Zhou, Zhongyi, et al.
Published: (2023)
by: Zhou, Zhongyi, et al.
Published: (2023)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
by: Chai, Yuxiang, et al.
Published: (2024)
by: Chai, Yuxiang, et al.
Published: (2024)
UFO: A UI-Focused Agent for Windows OS Interaction
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
Generative UI: LLMs are Effective UI Generators
by: Leviathan, Yaniv, et al.
Published: (2026)
by: Leviathan, Yaniv, et al.
Published: (2026)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
by: Luera, Reuben A., et al.
Published: (2025)
by: Luera, Reuben A., et al.
Published: (2025)
On AI-Inspired UI-Design
by: Wei, Jialiang, et al.
Published: (2024)
by: Wei, Jialiang, et al.
Published: (2024)
TinyClick: Single-Turn Agent for Empowering GUI Automation
by: Pawlowski, Pawel, et al.
Published: (2024)
by: Pawlowski, Pawel, et al.
Published: (2024)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
Acquiring Grounded Representations of Words with Situated Interactive Instruction
by: Mohan, Shiwali, et al.
Published: (2025)
by: Mohan, Shiwali, et al.
Published: (2025)
Similar Items
-
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024) -
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025) -
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024) -
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025) -
MAIC-UI: Making Interactive Courseware with Generative UI
by: Tu, Shangqing, et al.
Published: (2026)