AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
Fuente:
arXiv
Saved in:
| Main Authors: | Anand, Neeraj, Jain, Rishabh, Patnaik, Sohan, Krishnamurthy, Balaji, Sarkar, Mausoom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
by: SR, Nikitha, et al.
Published: (2025)
by: SR, Nikitha, et al.
Published: (2025)
EOPose : Exemplar-based object reposing using Generalized Pose Correspondences
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
by: Gopal, Varun, et al.
Published: (2026)
by: Gopal, Varun, et al.
Published: (2026)
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
by: SR, Nikitha, et al.
Published: (2024)
by: SR, Nikitha, et al.
Published: (2024)
UCATSC: Uncertainty-Aware Constrained Traffic Signal Control Under Vision-Based Partial Observability
by: Bodagala, Jayawant, et al.
Published: (2026)
by: Bodagala, Jayawant, et al.
Published: (2026)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
by: Xu, Zhou, et al.
Published: (2026)
by: Xu, Zhou, et al.
Published: (2026)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
Adaptive Deep Iris Feature Extractor at Arbitrary Resolutions
by: Shoji, Yuho, et al.
Published: (2024)
by: Shoji, Yuho, et al.
Published: (2024)
DiffSign: AI-Assisted Generation of Customizable Sign Language Videos With Enhanced Realism
by: Krishnamurthy, Sudha, et al.
Published: (2024)
by: Krishnamurthy, Sudha, et al.
Published: (2024)
3D Reconstruction of Protein Structures from Multi-view AFM Images using Neural Radiance Fields (NeRFs)
by: Rade, Jaydeep, et al.
Published: (2024)
by: Rade, Jaydeep, et al.
Published: (2024)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
by: Shao, Rui, et al.
Published: (2026)
by: Shao, Rui, et al.
Published: (2026)
Early Fusion of Features for Semantic Segmentation
by: Gupta, Anupam, et al.
Published: (2024)
by: Gupta, Anupam, et al.
Published: (2024)
Measuring and Improving Persuasiveness of Large Language Models
by: Singh, Somesh, et al.
Published: (2024)
by: Singh, Somesh, et al.
Published: (2024)
CardioSAM: Topology-Aware Decoder Design for High-Precision Cardiac MRI Segmentation
by: Jain, Ujjwal
Published: (2026)
by: Jain, Ujjwal
Published: (2026)
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
by: Jadhav, Avadhoot, et al.
Published: (2025)
by: Jadhav, Avadhoot, et al.
Published: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
by: Srivastava, Ashutosh, et al.
Published: (2024)
by: Srivastava, Ashutosh, et al.
Published: (2024)
CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
by: Anand, Neeraj, et al.
Published: (2026)
by: Anand, Neeraj, et al.
Published: (2026)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
Depth-Aware Super-Resolution via Distance-Adaptive Variational Formulation
by: Guo, Tianhao, et al.
Published: (2025)
by: Guo, Tianhao, et al.
Published: (2025)
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026)
by: Zhao, Zhongyin, et al.
Published: (2026)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
by: Yao, Ruilin, et al.
Published: (2026)
by: Yao, Ruilin, et al.
Published: (2026)
Stro-VIGRU: Defining the Vision Recurrent-Based Baseline Model for Brain Stroke Classification
by: Das, Subhajeet, et al.
Published: (2025)
by: Das, Subhajeet, et al.
Published: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
DOOMGAN:High-Fidelity Dynamic Identity Obfuscation Ocular Generative Morphing
by: Krishnamurthy, Bharath, et al.
Published: (2025)
by: Krishnamurthy, Bharath, et al.
Published: (2025)
Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation
by: Gangwar, Neeraj, et al.
Published: (2025)
by: Gangwar, Neeraj, et al.
Published: (2025)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
by: Liu, Enqi, et al.
Published: (2026)
by: Liu, Enqi, et al.
Published: (2026)
Real-Time Bundle Adjustment for Ultra-High-Resolution UAV Imagery Using Adaptive Patch-Based Feature Tracking
by: Iz, Selim Ahmet, et al.
Published: (2025)
by: Iz, Selim Ahmet, et al.
Published: (2025)
Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment
by: Wang, Haobo, et al.
Published: (2026)
by: Wang, Haobo, et al.
Published: (2026)
Dynamic Motion Synthesis: Masked Audio-Text Conditioned Spatio-Temporal Transformers
by: Anisetty, Sohan, et al.
Published: (2024)
by: Anisetty, Sohan, et al.
Published: (2024)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
by: Shakeel, Rozain, et al.
Published: (2026)
by: Shakeel, Rozain, et al.
Published: (2026)
Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation Transformer
by: Ai, Yuang, et al.
Published: (2023)
by: Ai, Yuang, et al.
Published: (2023)
GeoSynth: Contextually-Aware High-Resolution Satellite Image Synthesis
by: Sastry, Srikumar, et al.
Published: (2024)
by: Sastry, Srikumar, et al.
Published: (2024)
GramSR: Visual Feature Conditioning for Diffusion-Based Super-Resolution
by: D'Oronzio, Fabio, et al.
Published: (2026)
by: D'Oronzio, Fabio, et al.
Published: (2026)
Streamlined Global and Local Features Combinator (SGLC) for High Resolution Image Dehazing
by: Benjdira, Bilel, et al.
Published: (2023)
by: Benjdira, Bilel, et al.
Published: (2023)
HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection
by: Zhou, Han, et al.
Published: (2026)
by: Zhou, Han, et al.
Published: (2026)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
Similar Items
-
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
by: Patnaik, Sohan, et al.
Published: (2025) -
HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
by: SR, Nikitha, et al.
Published: (2025) -
EOPose : Exemplar-based object reposing using Generalized Pose Correspondences
by: Mehrotra, Sarthak, et al.
Published: (2025) -
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
by: Gopal, Varun, et al.
Published: (2026) -
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
by: SR, Nikitha, et al.
Published: (2024)