iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mehrotra, Sarthak, Rebbapragada, Sairam V C, Bonthu, Mani Hemanth Reddy, Balasubramanian, Vineeth N |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024)
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024)
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026)
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
von: VCR, Sairam, et al.
Veröffentlicht: (2025)
von: VCR, Sairam, et al.
Veröffentlicht: (2025)
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2024)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2024)
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
UItron: Foundational GUI Agent with Advanced Perception and Planning
von: Zeng, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zeng, Zhixiong, et al.
Veröffentlicht: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach
von: Rathore, Vaibhav, et al.
Veröffentlicht: (2025)
von: Rathore, Vaibhav, et al.
Veröffentlicht: (2025)
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
von: Devalapally, Arnav, et al.
Veröffentlicht: (2026)
von: Devalapally, Arnav, et al.
Veröffentlicht: (2026)
Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance
von: Santra, Sanchayan, et al.
Veröffentlicht: (2025)
von: Santra, Sanchayan, et al.
Veröffentlicht: (2025)
REJEPA: A Novel Joint-Embedding Predictive Architecture for Efficient Remote Sensing Image Retrieval
von: Choudhury, Shabnam, et al.
Veröffentlicht: (2025)
von: Choudhury, Shabnam, et al.
Veröffentlicht: (2025)
LogicCBMs: Logic-Enhanced Concept-Based Learning
von: Vemuri, Deepika SN, et al.
Veröffentlicht: (2025)
von: Vemuri, Deepika SN, et al.
Veröffentlicht: (2025)
Advancing Ante-Hoc Explainable Models through Generative Adversarial Networks
von: Garg, Tanmay, et al.
Veröffentlicht: (2024)
von: Garg, Tanmay, et al.
Veröffentlicht: (2024)
Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
von: Khindkar, Vaishnavi, et al.
Veröffentlicht: (2024)
von: Khindkar, Vaishnavi, et al.
Veröffentlicht: (2024)
Open-Set Object Detection By Aligning Known Class Representations
von: Sarkar, Hiran, et al.
Veröffentlicht: (2024)
von: Sarkar, Hiran, et al.
Veröffentlicht: (2024)
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2024)
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2024)
Understanding Task Transfer in Vision-Language Models
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
von: Schusterbauer, Johannes, et al.
Veröffentlicht: (2026)
von: Schusterbauer, Johannes, et al.
Veröffentlicht: (2026)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
EOPose : Exemplar-based object reposing using Generalized Pose Correspondences
von: Mehrotra, Sarthak, et al.
Veröffentlicht: (2025)
von: Mehrotra, Sarthak, et al.
Veröffentlicht: (2025)
VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
von: Kumar, Puneet, et al.
Veröffentlicht: (2022)
von: Kumar, Puneet, et al.
Veröffentlicht: (2022)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
von: Zeng, Haijin, et al.
Veröffentlicht: (2025)
von: Zeng, Haijin, et al.
Veröffentlicht: (2025)
SHIFT: Steering Hidden Intermediates in Flow Transformers
von: Konovalova, Nina, et al.
Veröffentlicht: (2026)
von: Konovalova, Nina, et al.
Veröffentlicht: (2026)
AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
von: Zhang, Ruifei, et al.
Veröffentlicht: (2025)
von: Zhang, Ruifei, et al.
Veröffentlicht: (2025)
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
Continual GUI Agents
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
Enhancing Visual Place Recognition via Fast and Slow Adaptive Biasing in Event Cameras
von: Nair, Gokul B., et al.
Veröffentlicht: (2024)
von: Nair, Gokul B., et al.
Veröffentlicht: (2024)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
CogAgent: A Visual Language Model for GUI Agents
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
von: Sinha, Rohit, et al.
Veröffentlicht: (2026)
von: Sinha, Rohit, et al.
Veröffentlicht: (2026)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
FaSTA$^*$: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing
von: Gupta, Advait, et al.
Veröffentlicht: (2025)
von: Gupta, Advait, et al.
Veröffentlicht: (2025)
Fiducial Focus Augmentation for Facial Landmark Detection
von: Kar, Purbayan, et al.
Veröffentlicht: (2024)
von: Kar, Purbayan, et al.
Veröffentlicht: (2024)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
von: Rebbapragada, Sairam VC, et al.
Veröffentlicht: (2024) -
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
von: Pathak, Harsharaj, et al.
Veröffentlicht: (2026) -
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025) -
Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
von: VCR, Sairam, et al.
Veröffentlicht: (2025) -
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2024)