Falcon-UI: Understanding GUI Before Following User Instructions
Fuente:
arXiv
Guardado en:
| Autores principales: | Shen, Huawen, Liu, Chang, Li, Gengluo, Wang, Xinlong, Zhou, Yu, Ma, Can, Ji, Xiangyang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
por: Shen, Huawen, et al.
Publicado: (2024)
por: Shen, Huawen, et al.
Publicado: (2024)
Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
por: Li, Gengluo, et al.
Publicado: (2025)
por: Li, Gengluo, et al.
Publicado: (2025)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
por: Qin, Yujia, et al.
Publicado: (2025)
por: Qin, Yujia, et al.
Publicado: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
por: Liu, Xinyi, et al.
Publicado: (2025)
por: Liu, Xinyi, et al.
Publicado: (2025)
Resolving Sentiment Discrepancy for Multimodal Sentiment Detection via Semantics Completion and Decomposition
por: Wu, Daiqing, et al.
Publicado: (2024)
por: Wu, Daiqing, et al.
Publicado: (2024)
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
por: Zhang, Yan, et al.
Publicado: (2026)
por: Zhang, Yan, et al.
Publicado: (2026)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
por: Tang, Fei, et al.
Publicado: (2026)
por: Tang, Fei, et al.
Publicado: (2026)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
por: Lian, Shuquan, et al.
Publicado: (2025)
por: Lian, Shuquan, et al.
Publicado: (2025)
Enhancing and Assessing Instruction-Following with Fine-Grained Instruction Variants
por: Yang, Jiuding, et al.
Publicado: (2024)
por: Yang, Jiuding, et al.
Publicado: (2024)
Aria-UI: Visual Grounding for GUI Instructions
por: Yang, Yuhao, et al.
Publicado: (2024)
por: Yang, Yuhao, et al.
Publicado: (2024)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
por: Wang, Haoming, et al.
Publicado: (2025)
por: Wang, Haoming, et al.
Publicado: (2025)
Infer Human's Intentions Before Following Natural Language Instructions
por: Wan, Yanming, et al.
Publicado: (2024)
por: Wan, Yanming, et al.
Publicado: (2024)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
por: Jing, Hongyi, et al.
Publicado: (2025)
por: Jing, Hongyi, et al.
Publicado: (2025)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
por: Li, Zhangheng, et al.
Publicado: (2024)
por: Li, Zhangheng, et al.
Publicado: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
por: Xiao, Han, et al.
Publicado: (2025)
por: Xiao, Han, et al.
Publicado: (2025)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
por: Zang, Yuan, et al.
Publicado: (2025)
por: Zang, Yuan, et al.
Publicado: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
por: Chen, Dongping, et al.
Publicado: (2024)
por: Chen, Dongping, et al.
Publicado: (2024)
MAIC-UI: Making Interactive Courseware with Generative UI
por: Tu, Shangqing, et al.
Publicado: (2026)
por: Tu, Shangqing, et al.
Publicado: (2026)
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
por: Jia, Qi, et al.
Publicado: (2025)
por: Jia, Qi, et al.
Publicado: (2025)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
por: You, Keen, et al.
Publicado: (2024)
por: You, Keen, et al.
Publicado: (2024)
MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation
por: Liu, Yile, et al.
Publicado: (2025)
por: Liu, Yile, et al.
Publicado: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
por: Nayak, Shravan, et al.
Publicado: (2025)
por: Nayak, Shravan, et al.
Publicado: (2025)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
por: Yang, Zhen, et al.
Publicado: (2025)
por: Yang, Zhen, et al.
Publicado: (2025)
Incubating Text Classifiers Following User Instruction with Nothing but LLM
por: Peng, Letian, et al.
Publicado: (2024)
por: Peng, Letian, et al.
Publicado: (2024)
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
por: Han, Wenkang, et al.
Publicado: (2025)
por: Han, Wenkang, et al.
Publicado: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
por: Wu, Xuansheng, et al.
Publicado: (2023)
por: Wu, Xuansheng, et al.
Publicado: (2023)
Identifying User Goals from UI Trajectories
por: Berkovitch, Omri, et al.
Publicado: (2024)
por: Berkovitch, Omri, et al.
Publicado: (2024)
Morae: Proactively Pausing UI Agents for User Choices
por: Peng, Yi-Hao, et al.
Publicado: (2025)
por: Peng, Yi-Hao, et al.
Publicado: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
por: Zhou, Yuqi, et al.
Publicado: (2025)
por: Zhou, Yuqi, et al.
Publicado: (2025)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
por: Chen, Liangyu, et al.
Publicado: (2025)
por: Chen, Liangyu, et al.
Publicado: (2025)
UltraIF: Advancing Instruction Following from the Wild
por: An, Kaikai, et al.
Publicado: (2025)
por: An, Kaikai, et al.
Publicado: (2025)
Instruction Following without Instruction Tuning
por: Hewitt, John, et al.
Publicado: (2024)
por: Hewitt, John, et al.
Publicado: (2024)
Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
por: Jeon, Jaehyun, et al.
Publicado: (2025)
por: Jeon, Jaehyun, et al.
Publicado: (2025)
StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following
por: Li, Jinnan, et al.
Publicado: (2025)
por: Li, Jinnan, et al.
Publicado: (2025)
Instruction-Following Pruning for Large Language Models
por: Hou, Bairu, et al.
Publicado: (2025)
por: Hou, Bairu, et al.
Publicado: (2025)
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
por: Yang, Wenkui, et al.
Publicado: (2026)
por: Yang, Wenkui, et al.
Publicado: (2026)
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
por: Wu, Qinzhuo, et al.
Publicado: (2024)
por: Wu, Qinzhuo, et al.
Publicado: (2024)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
por: Liang, Chen, et al.
Publicado: (2026)
por: Liang, Chen, et al.
Publicado: (2026)
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
por: Li, Gengluo, et al.
Publicado: (2026)
por: Li, Gengluo, et al.
Publicado: (2026)
Ejemplares similares
-
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
por: Shen, Huawen, et al.
Publicado: (2024) -
Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
por: Li, Gengluo, et al.
Publicado: (2025) -
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
por: Qin, Yujia, et al.
Publicado: (2025) -
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
por: Liu, Xinyi, et al.
Publicado: (2025) -
Resolving Sentiment Discrepancy for Multimodal Sentiment Detection via Semantics Completion and Decomposition
por: Wu, Daiqing, et al.
Publicado: (2024)