AUTONODE: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
Fuente:
arXiv
Saved in:
| Main Authors: | Datta, Arkajit, Verma, Tushar, Chawla, Rajat, S, Mukunda N., Bhola, Ishaan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
by: Rahman, Abdur, et al.
Published: (2024)
by: Rahman, Abdur, et al.
Published: (2024)
Veagle: Advancements in Multimodal Representation Learning
by: Chawla, Rajat, et al.
Published: (2024)
by: Chawla, Rajat, et al.
Published: (2024)
GUIDE: Graphical User Interface Data for Execution
by: Chawla, Rajat, et al.
Published: (2024)
by: Chawla, Rajat, et al.
Published: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Longitudinal Boundary Sharpness Coefficient Slopes Predict Time to Alzheimer's Disease Conversion in Mild Cognitive Impairment: A Survival Analysis Using the ADNI Cohort
by: Cherukuri, Ishaan
Published: (2026)
by: Cherukuri, Ishaan
Published: (2026)
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
by: Yang, Chenyu, et al.
Published: (2025)
by: Yang, Chenyu, et al.
Published: (2025)
NeuroBridge: Bio-Inspired Self-Supervised EEG-to-Image Decoding via Cognitive Priors and Bidirectional Semantic Alignment
by: Zhang, Wenjiang, et al.
Published: (2025)
by: Zhang, Wenjiang, et al.
Published: (2025)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
by: Domkundwar, Ishaan, et al.
Published: (2024)
by: Domkundwar, Ishaan, et al.
Published: (2024)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
by: Sun, Yuchen, et al.
Published: (2025)
by: Sun, Yuchen, et al.
Published: (2025)
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
by: Wang, Wanfu, et al.
Published: (2025)
by: Wang, Wanfu, et al.
Published: (2025)
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
by: Zhang, Yan, et al.
Published: (2026)
by: Zhang, Yan, et al.
Published: (2026)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training
by: Chen, Yuxi, et al.
Published: (2026)
by: Chen, Yuxi, et al.
Published: (2026)
GPA: Learning GUI Process Automation from Demonstrations
by: Zhao, Zirui, et al.
Published: (2026)
by: Zhao, Zirui, et al.
Published: (2026)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
Interactive Video Generation via Domain Adaptation
by: Rawal, Ishaan, et al.
Published: (2025)
by: Rawal, Ishaan, et al.
Published: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
SOAR: Advancements in Small Body Object Detection for Aerial Imagery Using State Space Models and Programmable Gradients
by: Verma, Tushar, et al.
Published: (2024)
by: Verma, Tushar, et al.
Published: (2024)
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
by: Lin, Juekai, et al.
Published: (2026)
by: Lin, Juekai, et al.
Published: (2026)
Graph4GUI: Graph Neural Networks for Representing Graphical User Interfaces
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
by: Rawal, Ishaan, et al.
Published: (2026)
by: Rawal, Ishaan, et al.
Published: (2026)
Di3PO - Diptych Diffusion DPO for Targeted Improvements in Image Generation
by: Reddy, Sanjana, et al.
Published: (2026)
by: Reddy, Sanjana, et al.
Published: (2026)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants
by: Hu, Junan, et al.
Published: (2026)
by: Hu, Junan, et al.
Published: (2026)
BAMI: Training-Free Bias Mitigation in GUI Grounding
by: Zhang, Borui, et al.
Published: (2026)
by: Zhang, Borui, et al.
Published: (2026)
GEBench: Benchmarking Image Generation Models as GUI Environments
by: Li, Haodong, et al.
Published: (2026)
by: Li, Haodong, et al.
Published: (2026)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
Implicit Deformable Medical Image Registration with Learnable Kernels
by: Fogarollo, Stefano, et al.
Published: (2025)
by: Fogarollo, Stefano, et al.
Published: (2025)
Asynchronous Perception Machine For Efficient Test-Time-Training
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025)
by: Qin, Yujia, et al.
Published: (2025)
NeuroAPS-Net: Neuro-Anatomically Aware Point Cloud Representation for Efficient Alzheimer's Disease Classification
by: Islam, Towhidul, et al.
Published: (2026)
by: Islam, Towhidul, et al.
Published: (2026)
Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation
by: Gao, Xinzge, et al.
Published: (2025)
by: Gao, Xinzge, et al.
Published: (2025)
Towards Neuro-Symbolic Video Understanding
by: Choi, Minkyu, et al.
Published: (2024)
by: Choi, Minkyu, et al.
Published: (2024)
UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks
by: Nguyen, Jason, et al.
Published: (2026)
by: Nguyen, Jason, et al.
Published: (2026)
Similar Items
-
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
by: Rahman, Abdur, et al.
Published: (2024) -
Veagle: Advancements in Multimodal Representation Learning
by: Chawla, Rajat, et al.
Published: (2024) -
GUIDE: Graphical User Interface Data for Execution
by: Chawla, Rajat, et al.
Published: (2024) -
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024) -
Longitudinal Boundary Sharpness Coefficient Slopes Predict Time to Alzheimer's Disease Conversion in Mild Cognitive Impairment: A Survival Analysis Using the ADNI Cohort
by: Cherukuri, Ishaan
Published: (2026)