Salvato in:
| Autori principali: | Zhou, Xurui, Chen, Gongwei, Xie, Yuquan, Li, Zaijing, Zhou, Kaiwen, Wang, Shuai, Yang, Shuo, Tian, Zhuotao, Shao, Rui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2512.01763 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Less is More: Empowering GUI Agent with Context-Aware Simplification
di: Chen, Gongwei, et al.
Pubblicazione: (2025)
di: Chen, Gongwei, et al.
Pubblicazione: (2025)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
di: Shao, Rui, et al.
Pubblicazione: (2026)
di: Shao, Rui, et al.
Pubblicazione: (2026)
Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
di: Xie, Yuquan, et al.
Pubblicazione: (2025)
di: Xie, Yuquan, et al.
Pubblicazione: (2025)
Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy
di: Li, Zaijing, et al.
Pubblicazione: (2025)
di: Li, Zaijing, et al.
Pubblicazione: (2025)
GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
di: Xie, Bin, et al.
Pubblicazione: (2025)
di: Xie, Bin, et al.
Pubblicazione: (2025)
Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
di: Li, Zaijing, et al.
Pubblicazione: (2024)
di: Li, Zaijing, et al.
Pubblicazione: (2024)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
di: Zhou, Yuqi, et al.
Pubblicazione: (2025)
di: Zhou, Yuqi, et al.
Pubblicazione: (2025)
Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimization
di: Li, Zaijing, et al.
Pubblicazione: (2025)
di: Li, Zaijing, et al.
Pubblicazione: (2025)
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
di: Lyu, Yibo, et al.
Pubblicazione: (2026)
di: Lyu, Yibo, et al.
Pubblicazione: (2026)
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
di: Shi, Haoxiang, et al.
Pubblicazione: (2025)
di: Shi, Haoxiang, et al.
Pubblicazione: (2025)
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
di: Li, Zaijing, et al.
Pubblicazione: (2026)
di: Li, Zaijing, et al.
Pubblicazione: (2026)
History-Aware Reasoning for GUI Agents
di: Wang, Ziwei, et al.
Pubblicazione: (2025)
di: Wang, Ziwei, et al.
Pubblicazione: (2025)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
di: Li, Wei, et al.
Pubblicazione: (2025)
di: Li, Wei, et al.
Pubblicazione: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
di: Xie, Yiping, et al.
Pubblicazione: (2026)
di: Xie, Yiping, et al.
Pubblicazione: (2026)
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
di: Zhang, Renshan, et al.
Pubblicazione: (2025)
di: Zhang, Renshan, et al.
Pubblicazione: (2025)
Enhancing Emotional Generation Capability of Large Language Models via Emotional Chain-of-Thought
di: Li, Zaijing, et al.
Pubblicazione: (2024)
di: Li, Zaijing, et al.
Pubblicazione: (2024)
HCQA @ Ego4D EgoSchema Challenge 2024
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
ObjectNLQ @ Ego4D Episodic Memory Challenge 2024
di: Feng, Yisen, et al.
Pubblicazione: (2024)
di: Feng, Yisen, et al.
Pubblicazione: (2024)
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
di: Lu, Fanbin, et al.
Pubblicazione: (2025)
di: Lu, Fanbin, et al.
Pubblicazione: (2025)
MobileFlow: A Multimodal LLM For Mobile GUI Agent
di: Nong, Songqin, et al.
Pubblicazione: (2024)
di: Nong, Songqin, et al.
Pubblicazione: (2024)
Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents
di: Song, Yurun, et al.
Pubblicazione: (2026)
di: Song, Yurun, et al.
Pubblicazione: (2026)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
di: Wang, Xuehui, et al.
Pubblicazione: (2025)
di: Wang, Xuehui, et al.
Pubblicazione: (2025)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
di: Xu, Zhou, et al.
Pubblicazione: (2026)
di: Xu, Zhou, et al.
Pubblicazione: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
POINTS-GUI-G: GUI-Grounding Journey
di: Zhao, Zhongyin, et al.
Pubblicazione: (2026)
di: Zhao, Zhongyin, et al.
Pubblicazione: (2026)
Continual GUI Agents
di: Liu, Ziwei, et al.
Pubblicazione: (2026)
di: Liu, Ziwei, et al.
Pubblicazione: (2026)
M$^2$-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining
di: Lv, Rui, et al.
Pubblicazione: (2026)
di: Lv, Rui, et al.
Pubblicazione: (2026)
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
di: Li, Yang, et al.
Pubblicazione: (2026)
di: Li, Yang, et al.
Pubblicazione: (2026)
AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech
di: Kang, Bin, et al.
Pubblicazione: (2026)
di: Kang, Bin, et al.
Pubblicazione: (2026)
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
di: Bu, Tianpeng, et al.
Pubblicazione: (2026)
di: Bu, Tianpeng, et al.
Pubblicazione: (2026)
MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
di: Zhou, Hanzhang, et al.
Pubblicazione: (2025)
di: Zhou, Hanzhang, et al.
Pubblicazione: (2025)
Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation
di: Shao, Tong, et al.
Pubblicazione: (2024)
di: Shao, Tong, et al.
Pubblicazione: (2024)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
di: Liu, Enqi, et al.
Pubblicazione: (2026)
di: Liu, Enqi, et al.
Pubblicazione: (2026)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
di: Shen, Leyang, et al.
Pubblicazione: (2024)
di: Shen, Leyang, et al.
Pubblicazione: (2024)
HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection
di: Zhou, Shuchang, et al.
Pubblicazione: (2026)
di: Zhou, Shuchang, et al.
Pubblicazione: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
di: Ye, Xianhang, et al.
Pubblicazione: (2025)
di: Ye, Xianhang, et al.
Pubblicazione: (2025)
CogAgent: A Visual Language Model for GUI Agents
di: Hong, Wenyi, et al.
Pubblicazione: (2023)
di: Hong, Wenyi, et al.
Pubblicazione: (2023)
Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization
di: Shao, Tong, et al.
Pubblicazione: (2025)
di: Shao, Tong, et al.
Pubblicazione: (2025)
MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples
di: Li, Xurui, et al.
Pubblicazione: (2025)
di: Li, Xurui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Less is More: Empowering GUI Agent with Context-Aware Simplification
di: Chen, Gongwei, et al.
Pubblicazione: (2025) -
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
di: Shao, Rui, et al.
Pubblicazione: (2026) -
Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
di: Xie, Yuquan, et al.
Pubblicazione: (2025) -
Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy
di: Li, Zaijing, et al.
Pubblicazione: (2025) -
GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
di: Xie, Bin, et al.
Pubblicazione: (2025)