V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Rahman, Abdur, Chawla, Rajat, Kumar, Muskaan, Datta, Arkajit, Jha, Adarsh, NS, Mukunda, Bhola, Ishaan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GUIDE: Graphical User Interface Data for Execution
by: Chawla, Rajat, et al.
Published: (2024)
by: Chawla, Rajat, et al.
Published: (2024)
AUTONODE: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
by: Datta, Arkajit, et al.
Published: (2024)
by: Datta, Arkajit, et al.
Published: (2024)
Veagle: Advancements in Multimodal Representation Learning
by: Chawla, Rajat, et al.
Published: (2024)
by: Chawla, Rajat, et al.
Published: (2024)
SuperCoder2.0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer
by: Gautam, Anmol, et al.
Published: (2024)
by: Gautam, Anmol, et al.
Published: (2024)
Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
by: Bhola, Ishaan, et al.
Published: (2025)
by: Bhola, Ishaan, et al.
Published: (2025)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
by: Domkundwar, Ishaan, et al.
Published: (2024)
by: Domkundwar, Ishaan, et al.
Published: (2024)
Multimodal Fusion of Glucose Monitoring and Food Imagery for Caloric Content Prediction
by: Kumar, Adarsh
Published: (2025)
by: Kumar, Adarsh
Published: (2025)
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026)
by: Zhao, Zhongyin, et al.
Published: (2026)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
by: Zhou, Yuqi, et al.
Published: (2025)
by: Zhou, Yuqi, et al.
Published: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
by: Ma, Longhui, et al.
Published: (2026)
by: Ma, Longhui, et al.
Published: (2026)
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
by: Shakeel, Rozain, et al.
Published: (2026)
by: Shakeel, Rozain, et al.
Published: (2026)
MobileFlow: A Multimodal LLM For Mobile GUI Agent
by: Nong, Songqin, et al.
Published: (2024)
by: Nong, Songqin, et al.
Published: (2024)
WeCKD: Weakly-supervised Chained Distillation Network for Efficient Multimodal Medical Imaging
by: Rahman, Md. Abdur, et al.
Published: (2025)
by: Rahman, Md. Abdur, et al.
Published: (2025)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
by: Lian, Shuquan, et al.
Published: (2025)
by: Lian, Shuquan, et al.
Published: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning
by: Xu, Hai-Ming, et al.
Published: (2024)
by: Xu, Hai-Ming, et al.
Published: (2024)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
Interactive Video Generation via Domain Adaptation
by: Rawal, Ishaan, et al.
Published: (2025)
by: Rawal, Ishaan, et al.
Published: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
by: Zhang, Miaosen, et al.
Published: (2025)
by: Zhang, Miaosen, et al.
Published: (2025)
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
by: Li, Weiming, et al.
Published: (2025)
by: Li, Weiming, et al.
Published: (2025)
On the Robustness of GUI Grounding Models Against Image Attacks
by: Zhao, Haoren, et al.
Published: (2025)
by: Zhao, Haoren, et al.
Published: (2025)
MVP: Multiple View Prediction Improves GUI Grounding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
by: Zeng, Ziyun, et al.
Published: (2026)
by: Zeng, Ziyun, et al.
Published: (2026)
MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
by: Deria, Ankan, et al.
Published: (2026)
by: Deria, Ankan, et al.
Published: (2026)
LinMU: Multimodal Understanding Made Linear
by: Wang, Hongjie, et al.
Published: (2026)
by: Wang, Hongjie, et al.
Published: (2026)
Improved GUI Grounding via Iterative Narrowing
by: Nguyen, Anthony
Published: (2024)
by: Nguyen, Anthony
Published: (2024)
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
by: Li, Hongxin, et al.
Published: (2026)
by: Li, Hongxin, et al.
Published: (2026)
Punching Above Precision: Small Quantized Model Distillation with Learnable Regularizer
by: Rehman, Abdur, et al.
Published: (2025)
by: Rehman, Abdur, et al.
Published: (2025)
Similar Items
-
GUIDE: Graphical User Interface Data for Execution
by: Chawla, Rajat, et al.
Published: (2024) -
AUTONODE: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
by: Datta, Arkajit, et al.
Published: (2024) -
Veagle: Advancements in Multimodal Representation Learning
by: Chawla, Rajat, et al.
Published: (2024) -
SuperCoder2.0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer
by: Gautam, Anmol, et al.
Published: (2024) -
Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
by: Bhola, Ishaan, et al.
Published: (2025)