DocOS: Towards Proactive Document-Guided Actions in GUI Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jingjing, Huang, Ziye, Cheng, Zihao, Liu, Zeming, Wu, Jiahong, Guo, Yuhang, Chen, Kehai, Wang, Yunhong, Wang, Haifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ToolSpectrum : Towards Personalized Tool Utilization for Large Language Models
von: Cheng, Zihao, et al.
Veröffentlicht: (2025)
von: Cheng, Zihao, et al.
Veröffentlicht: (2025)
DocMEdit: Towards Document-Level Model Editing
von: Zeng, Li, et al.
Veröffentlicht: (2025)
von: Zeng, Li, et al.
Veröffentlicht: (2025)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
von: Wu, Zheng, et al.
Veröffentlicht: (2025)
von: Wu, Zheng, et al.
Veröffentlicht: (2025)
Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
von: Lu, Yuheng, et al.
Veröffentlicht: (2025)
von: Lu, Yuheng, et al.
Veröffentlicht: (2025)
Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
von: Li, Silin, et al.
Veröffentlicht: (2025)
von: Li, Silin, et al.
Veröffentlicht: (2025)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
CogDoc: Towards Unified thinking in Documents
von: Xu, Qixin, et al.
Veröffentlicht: (2025)
von: Xu, Qixin, et al.
Veröffentlicht: (2025)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
von: Xia, Hongfei, et al.
Veröffentlicht: (2025)
von: Xia, Hongfei, et al.
Veröffentlicht: (2025)
GUI-PRA: Process Reward Agent for GUI Tasks
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents
von: Chai, Yuxiang, et al.
Veröffentlicht: (2026)
von: Chai, Yuxiang, et al.
Veröffentlicht: (2026)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
von: Liu, Yibai, et al.
Veröffentlicht: (2025)
von: Liu, Yibai, et al.
Veröffentlicht: (2025)
Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
Faithful Mobile GUI Agents with Guided Advantage Estimator
von: Hu, Haowen, et al.
Veröffentlicht: (2026)
von: Hu, Haowen, et al.
Veröffentlicht: (2026)
Deepfake Detection via Knowledge Injection
von: Li, Tonghui, et al.
Veröffentlicht: (2025)
von: Li, Tonghui, et al.
Veröffentlicht: (2025)
Agentic Reward Modeling: Verifying GUI Agent via Online Proactive Interaction
von: Cui, Chaoqun, et al.
Veröffentlicht: (2026)
von: Cui, Chaoqun, et al.
Veröffentlicht: (2026)
TCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine
von: Cheng, Zihao, et al.
Veröffentlicht: (2025)
von: Cheng, Zihao, et al.
Veröffentlicht: (2025)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
Deterministic Reversible Data Augmentation for Neural Machine Translation
von: Yao, Jiashu, et al.
Veröffentlicht: (2024)
von: Yao, Jiashu, et al.
Veröffentlicht: (2024)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
Towards Trustworthy GUI Agents: A Survey
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
FAME: Towards Factual Multi-Task Model Editing
von: Zeng, Li, et al.
Veröffentlicht: (2024)
von: Zeng, Li, et al.
Veröffentlicht: (2024)
RETAIL: Towards Real-world Travel Planning for Large Language Models
von: Deng, Bin, et al.
Veröffentlicht: (2025)
von: Deng, Bin, et al.
Veröffentlicht: (2025)
OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
Generalizable AI-Generated Image Detection Based on Fractal Self-Similarity in the Spectrum
von: Xiao, Shengpeng, et al.
Veröffentlicht: (2025)
von: Xiao, Shengpeng, et al.
Veröffentlicht: (2025)
OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
von: Li, Zehao, et al.
Veröffentlicht: (2026)
von: Li, Zehao, et al.
Veröffentlicht: (2026)
PRIM: Towards Practical In-Image Multilingual Machine Translation
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
Exploring In-Image Machine Translation with Real-World Background
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
DocDancer: Towards Agentic Document-Grounded Information Seeking
von: Zhang, Qintong, et al.
Veröffentlicht: (2026)
von: Zhang, Qintong, et al.
Veröffentlicht: (2026)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2025)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2025)
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants
von: Hu, Junan, et al.
Veröffentlicht: (2026)
von: Hu, Junan, et al.
Veröffentlicht: (2026)
DocAgent: A Multi-Agent System for Automated Code Documentation Generation
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ToolSpectrum : Towards Personalized Tool Utilization for Large Language Models
von: Cheng, Zihao, et al.
Veröffentlicht: (2025) -
DocMEdit: Towards Document-Level Model Editing
von: Zeng, Li, et al.
Veröffentlicht: (2025) -
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
von: Liu, Jingjing, et al.
Veröffentlicht: (2025) -
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
von: Wu, Zheng, et al.
Veröffentlicht: (2025) -
Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)