InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yuhang, Li, Pengxiang, Wei, Zishu, Xie, Congkai, Hu, Xueyu, Xu, Xinchen, Zhang, Shengyu, Han, Xiaotian, Yang, Hongxia, Wu, Fei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios
di: Yu, Chenglin, et al.
Pubblicazione: (2025)
di: Yu, Chenglin, et al.
Pubblicazione: (2025)
InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
di: Wu, Zhiyong, et al.
Pubblicazione: (2024)
di: Wu, Zhiyong, et al.
Pubblicazione: (2024)
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
di: Liu, Guangyi, et al.
Pubblicazione: (2025)
di: Liu, Guangyi, et al.
Pubblicazione: (2025)
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
di: Liu, Guangyi, et al.
Pubblicazione: (2025)
di: Liu, Guangyi, et al.
Pubblicazione: (2025)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
di: Liao, Zeyi, et al.
Pubblicazione: (2025)
di: Liao, Zeyi, et al.
Pubblicazione: (2025)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
di: Qin, Yujia, et al.
Pubblicazione: (2025)
di: Qin, Yujia, et al.
Pubblicazione: (2025)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
di: Nong, Songqin, et al.
Pubblicazione: (2025)
di: Nong, Songqin, et al.
Pubblicazione: (2025)
GUI Agents: A Survey
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
History-Aware Reasoning for GUI Agents
di: Wang, Ziwei, et al.
Pubblicazione: (2025)
di: Wang, Ziwei, et al.
Pubblicazione: (2025)
Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition
di: Zhou, Zishu, et al.
Pubblicazione: (2026)
di: Zhou, Zishu, et al.
Pubblicazione: (2026)
WinClick: GUI Grounding with Multimodal Large Language Models
di: Hui, Zheng, et al.
Pubblicazione: (2025)
di: Hui, Zheng, et al.
Pubblicazione: (2025)
Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
di: Zhang, Shuning, et al.
Pubblicazione: (2025)
di: Zhang, Shuning, et al.
Pubblicazione: (2025)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
di: Cheng, Kanzhi, et al.
Pubblicazione: (2024)
di: Cheng, Kanzhi, et al.
Pubblicazione: (2024)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
di: Henry, Felix, et al.
Pubblicazione: (2026)
di: Henry, Felix, et al.
Pubblicazione: (2026)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
di: Gebreegziabher, Simret Araya, et al.
Pubblicazione: (2026)
di: Gebreegziabher, Simret Araya, et al.
Pubblicazione: (2026)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
di: Wu, Zongru, et al.
Pubblicazione: (2025)
di: Wu, Zongru, et al.
Pubblicazione: (2025)
Exploring the Feasibility of Multimodal Chatbot AI as Copilot in Pathology Diagnostics: Generalist Model's Pitfall
di: Liu, Mianxin, et al.
Pubblicazione: (2024)
di: Liu, Mianxin, et al.
Pubblicazione: (2024)
Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation
di: Hao, Pan, et al.
Pubblicazione: (2026)
di: Hao, Pan, et al.
Pubblicazione: (2026)
API Agents vs. GUI Agents: Divergence and Convergence
di: Zhang, Chaoyun, et al.
Pubblicazione: (2025)
di: Zhang, Chaoyun, et al.
Pubblicazione: (2025)
GraphPilot: GUI Task Automation with One-Step LLM Reasoning Powered by Knowledge Graph
di: Yu, Mingxian, et al.
Pubblicazione: (2026)
di: Yu, Mingxian, et al.
Pubblicazione: (2026)
Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI
di: Mothilal, Ramaravind Kommiya, et al.
Pubblicazione: (2025)
di: Mothilal, Ramaravind Kommiya, et al.
Pubblicazione: (2025)
HealthPrism: A Visual Analytics System for Exploring Children's Physical and Mental Health Profiles with Multimodal Data
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
di: Tang, Liujian, et al.
Pubblicazione: (2025)
di: Tang, Liujian, et al.
Pubblicazione: (2025)
Enhancing Deliberativeness: Evaluating the Impact of Multimodal Reflection Nudges
di: Yeo, ShunYi, et al.
Pubblicazione: (2025)
di: Yeo, ShunYi, et al.
Pubblicazione: (2025)
GUI Agents with Foundation Models: A Comprehensive Survey
di: Wang, Shuai, et al.
Pubblicazione: (2024)
di: Wang, Shuai, et al.
Pubblicazione: (2024)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
di: Li, Hongxin, et al.
Pubblicazione: (2025)
di: Li, Hongxin, et al.
Pubblicazione: (2025)
Supporting Reflection and Forward-Looking Reasoning With Data-Driven Questions
di: Fischer, Simon WS, et al.
Pubblicazione: (2026)
di: Fischer, Simon WS, et al.
Pubblicazione: (2026)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
di: Cheng, Pengzhou, et al.
Pubblicazione: (2025)
di: Cheng, Pengzhou, et al.
Pubblicazione: (2025)
TinyClick: Single-Turn Agent for Empowering GUI Automation
di: Pawlowski, Pawel, et al.
Pubblicazione: (2024)
di: Pawlowski, Pawel, et al.
Pubblicazione: (2024)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
di: Chai, Yuxiang, et al.
Pubblicazione: (2024)
di: Chai, Yuxiang, et al.
Pubblicazione: (2024)
Large Language Model-Brained GUI Agents: A Survey
di: Zhang, Chaoyun, et al.
Pubblicazione: (2024)
di: Zhang, Chaoyun, et al.
Pubblicazione: (2024)
SpiritSight Agent: Advanced GUI Agent with One Look
di: Huang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Huang, Zhiyuan, et al.
Pubblicazione: (2025)
InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning
di: Xie, Congkai, et al.
Pubblicazione: (2025)
di: Xie, Congkai, et al.
Pubblicazione: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
di: Kapoor, Raghav, et al.
Pubblicazione: (2024)
di: Kapoor, Raghav, et al.
Pubblicazione: (2024)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
di: Lee, Jungjae, et al.
Pubblicazione: (2025)
di: Lee, Jungjae, et al.
Pubblicazione: (2025)
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
di: Gao, Longxi, et al.
Pubblicazione: (2024)
di: Gao, Longxi, et al.
Pubblicazione: (2024)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
di: Tang, Jingyu, et al.
Pubblicazione: (2025)
di: Tang, Jingyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
di: Liu, Yuhang, et al.
Pubblicazione: (2025) -
InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios
di: Yu, Chenglin, et al.
Pubblicazione: (2025) -
InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
di: Liu, Yuhang, et al.
Pubblicazione: (2025) -
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
di: Wu, Zhiyong, et al.
Pubblicazione: (2024) -
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
di: Liu, Guangyi, et al.
Pubblicazione: (2025)