Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Daiqiang, Pan, Zihao, Zhang, Zeyu, Chen, Ronghao, Wang, Huacan, Chen, Honggang, Jiang, Haiyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
von: Xu, Zhou, et al.
Veröffentlicht: (2026)
von: Xu, Zhou, et al.
Veröffentlicht: (2026)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
von: Takezoe, Rinyoichi, et al.
Veröffentlicht: (2026)
von: Takezoe, Rinyoichi, et al.
Veröffentlicht: (2026)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
von: Sha, Lin, et al.
Veröffentlicht: (2026)
von: Sha, Lin, et al.
Veröffentlicht: (2026)
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
von: Li, Yuankai, et al.
Veröffentlicht: (2026)
von: Li, Yuankai, et al.
Veröffentlicht: (2026)
Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning
von: Ma, Pengtao, et al.
Veröffentlicht: (2026)
von: Ma, Pengtao, et al.
Veröffentlicht: (2026)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
von: Tan, Zhentao, et al.
Veröffentlicht: (2024)
von: Tan, Zhentao, et al.
Veröffentlicht: (2024)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
von: Ding, Zijun, et al.
Veröffentlicht: (2025)
von: Ding, Zijun, et al.
Veröffentlicht: (2025)
HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
von: Zhu, Qihui, et al.
Veröffentlicht: (2026)
von: Zhu, Qihui, et al.
Veröffentlicht: (2026)
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
von: Ju, Shaobo, et al.
Veröffentlicht: (2026)
von: Ju, Shaobo, et al.
Veröffentlicht: (2026)
EEdit: Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing
von: Yan, Zexuan, et al.
Veröffentlicht: (2025)
von: Yan, Zexuan, et al.
Veröffentlicht: (2025)
VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
von: Deng, Jie, et al.
Veröffentlicht: (2026)
von: Deng, Jie, et al.
Veröffentlicht: (2026)
Focus-Scan-Refine: From Human Visual Perception to Efficient Visual Token Pruning
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
von: Zhuang, Jiedong, et al.
Veröffentlicht: (2024)
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
von: Li, Jiaqi, et al.
Veröffentlicht: (2026)
von: Li, Jiaqi, et al.
Veröffentlicht: (2026)
Select2Col: Leveraging Spatial-Temporal Importance of Semantic Information for Efficient Collaborative Perception
von: Liu, Yuntao, et al.
Veröffentlicht: (2023)
von: Liu, Yuntao, et al.
Veröffentlicht: (2023)
Training Noise Token Pruning
von: Rao, Mingxing, et al.
Veröffentlicht: (2024)
von: Rao, Mingxing, et al.
Veröffentlicht: (2024)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
CogAgent: A Visual Language Model for GUI Agents
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
Back to Fundamentals: Low-Level Visual Features Guided Progressive Token Pruning
von: Ouyang, Yuanbing, et al.
Veröffentlicht: (2025)
von: Ouyang, Yuanbing, et al.
Veröffentlicht: (2025)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph
von: Jiang, Yutao, et al.
Veröffentlicht: (2025)
von: Jiang, Yutao, et al.
Veröffentlicht: (2025)
OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations
von: Li, Zesheng, et al.
Veröffentlicht: (2026)
von: Li, Zesheng, et al.
Veröffentlicht: (2026)
CROP: Contextual Region-Oriented Visual Token Pruning
von: Guo, Jiawei, et al.
Veröffentlicht: (2025)
von: Guo, Jiawei, et al.
Veröffentlicht: (2025)
TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models
von: Cheng, Xinle, et al.
Veröffentlicht: (2025)
von: Cheng, Xinle, et al.
Veröffentlicht: (2025)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
TP-Spikformer: Token Pruned Spiking Transformer
von: Wei, Wenjie, et al.
Veröffentlicht: (2026)
von: Wei, Wenjie, et al.
Veröffentlicht: (2026)
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
von: He, Jialuo, et al.
Veröffentlicht: (2026)
von: He, Jialuo, et al.
Veröffentlicht: (2026)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
von: Ma, Jie, et al.
Veröffentlicht: (2026)
von: Ma, Jie, et al.
Veröffentlicht: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
von: Xu, Zhou, et al.
Veröffentlicht: (2026) -
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026) -
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
von: Takezoe, Rinyoichi, et al.
Veröffentlicht: (2026) -
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
von: Sha, Lin, et al.
Veröffentlicht: (2026) -
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
von: Li, Bo, et al.
Veröffentlicht: (2026)