Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Yurun, Yin, Jiong, Zhang, Rongjunchen, Harris, Ian G. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
von: Zhang, Chenxi, et al.
Veröffentlicht: (2026)
von: Zhang, Chenxi, et al.
Veröffentlicht: (2026)
HiconAgent: History Context-aware Policy Optimization for GUI Agents
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
Efficient Long-Horizon GUI Agents via Training-Free KV Cache Compression
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
von: Lu, Fanbin, et al.
Veröffentlicht: (2025)
von: Lu, Fanbin, et al.
Veröffentlicht: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
von: Zhu, Yuke, et al.
Veröffentlicht: (2024)
von: Zhu, Yuke, et al.
Veröffentlicht: (2024)
Efficient Masked Image Compression with Position-Indexed Self-Attention
von: Dai, Chengjie, et al.
Veröffentlicht: (2025)
von: Dai, Chengjie, et al.
Veröffentlicht: (2025)
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
von: Ma, Qiankun, et al.
Veröffentlicht: (2026)
von: Ma, Qiankun, et al.
Veröffentlicht: (2026)
CompressNAS : A Fast and Efficient Technique for Model Compression using Decomposition
von: Sah, Sudhakar, et al.
Veröffentlicht: (2025)
von: Sah, Sudhakar, et al.
Veröffentlicht: (2025)
MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
von: Zhang, Guohui, et al.
Veröffentlicht: (2025)
von: Zhang, Guohui, et al.
Veröffentlicht: (2025)
Efficient Large Multi-modal Models via Visual Context Compression
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Flow-Based Generative Modeling for Optimizing Sampling Policies in Compressed Sensing Applications
von: Pavelkin, Roman, et al.
Veröffentlicht: (2026)
von: Pavelkin, Roman, et al.
Veröffentlicht: (2026)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
AgentCompress: Task-Aware Compression for Affordable Large Language Model Agents
von: Taha, Zuhair Ahmed Khan, et al.
Veröffentlicht: (2026)
von: Taha, Zuhair Ahmed Khan, et al.
Veröffentlicht: (2026)
RobustSCI: Beyond Reconstruction to Restoration for Snapshot Compressive Imaging under Real-World Degradations
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
General Compression Framework for Efficient Transformer Object Tracking
von: Hong, Lingyi, et al.
Veröffentlicht: (2024)
von: Hong, Lingyi, et al.
Veröffentlicht: (2024)
A Multi-Stage Optimization Framework for Deploying Learned Image Compression on FPGAs
von: Fang, Jiaxun, et al.
Veröffentlicht: (2025)
von: Fang, Jiaxun, et al.
Veröffentlicht: (2025)
Efficient Reasoning via Thought Compression for Language Segmentation
von: Zhou, Qing, et al.
Veröffentlicht: (2026)
von: Zhou, Qing, et al.
Veröffentlicht: (2026)
SnapCap: Efficient Snapshot Compressive Video Captioning
von: Sun, Jianqiao, et al.
Veröffentlicht: (2024)
von: Sun, Jianqiao, et al.
Veröffentlicht: (2024)
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
von: Liu, Jiazheng, et al.
Veröffentlicht: (2025)
von: Liu, Jiazheng, et al.
Veröffentlicht: (2025)
Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
von: Afifi, Mahmoud, et al.
Veröffentlicht: (2025)
von: Afifi, Mahmoud, et al.
Veröffentlicht: (2025)
SCP: Spherical-Coordinate-based Learned Point Cloud Compression
von: Luo, Ao, et al.
Veröffentlicht: (2023)
von: Luo, Ao, et al.
Veröffentlicht: (2023)
DMOFC: Discrimination Metric-Optimized Feature Compression
von: Gao, Changsheng, et al.
Veröffentlicht: (2024)
von: Gao, Changsheng, et al.
Veröffentlicht: (2024)
Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers
von: Yang, Chaojie, et al.
Veröffentlicht: (2026)
von: Yang, Chaojie, et al.
Veröffentlicht: (2026)
Delta-SVD: Efficient Compression for Personalized Text-to-Image Models
von: Zhang, Tangyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Tangyuan, et al.
Veröffentlicht: (2025)
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
von: Yao, Ruilin, et al.
Veröffentlicht: (2026)
von: Yao, Ruilin, et al.
Veröffentlicht: (2026)
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
von: Mao, Weian, et al.
Veröffentlicht: (2026)
von: Mao, Weian, et al.
Veröffentlicht: (2026)
EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models
von: Lu, Hongyu, et al.
Veröffentlicht: (2026)
von: Lu, Hongyu, et al.
Veröffentlicht: (2026)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
CogAgent: A Visual Language Model for GUI Agents
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation
von: Liang, Guotao, et al.
Veröffentlicht: (2026)
von: Liang, Guotao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
von: Zhang, Chenxi, et al.
Veröffentlicht: (2026) -
HiconAgent: History Context-aware Policy Optimization for GUI Agents
von: Zhou, Xurui, et al.
Veröffentlicht: (2025) -
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
von: Yan, Haolong, et al.
Veröffentlicht: (2025) -
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025) -
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026)