Guardado en:
| Autores principales: | Lim, Jaehyuk, Lee, Bruce W. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2408.09111 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
por: Zhao, Henry Hengyuan, et al.
Publicado: (2025)
por: Zhao, Henry Hengyuan, et al.
Publicado: (2025)
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
por: Sun, Qiushi, et al.
Publicado: (2025)
por: Sun, Qiushi, et al.
Publicado: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
por: Kapoor, Raghav, et al.
Publicado: (2024)
por: Kapoor, Raghav, et al.
Publicado: (2024)
AIN: The Arabic INclusive Large Multimodal Model
por: Heakl, Ahmed, et al.
Publicado: (2025)
por: Heakl, Ahmed, et al.
Publicado: (2025)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
por: Lin, Ronghao, et al.
Publicado: (2025)
por: Lin, Ronghao, et al.
Publicado: (2025)
Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
por: Tsangko, Iosif, et al.
Publicado: (2025)
por: Tsangko, Iosif, et al.
Publicado: (2025)
Can Large Language Models Capture Video Game Engagement?
por: Melhart, David, et al.
Publicado: (2025)
por: Melhart, David, et al.
Publicado: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
por: Chen, Wentong, et al.
Publicado: (2024)
por: Chen, Wentong, et al.
Publicado: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
por: Chen, Baiyu, et al.
Publicado: (2026)
por: Chen, Baiyu, et al.
Publicado: (2026)
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
por: Huang, Ting-Hao 'Kenneth', et al.
Publicado: (2025)
por: Huang, Ting-Hao 'Kenneth', et al.
Publicado: (2025)
Code2World: A GUI World Model via Renderable Code Generation
por: Zheng, Yuhao, et al.
Publicado: (2026)
por: Zheng, Yuhao, et al.
Publicado: (2026)
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
por: Lopez-Cardona, Angela, et al.
Publicado: (2024)
por: Lopez-Cardona, Angela, et al.
Publicado: (2024)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
por: Tsangko, Iosif, et al.
Publicado: (2026)
por: Tsangko, Iosif, et al.
Publicado: (2026)
Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations
por: Pal, Ankit, et al.
Publicado: (2024)
por: Pal, Ankit, et al.
Publicado: (2024)
EEG-based Multimodal Representation Learning for Emotion Recognition
por: Yin, Kang, et al.
Publicado: (2024)
por: Yin, Kang, et al.
Publicado: (2024)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
por: Zhou, Shijie, et al.
Publicado: (2025)
por: Zhou, Shijie, et al.
Publicado: (2025)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
por: Xu, Kevin, et al.
Publicado: (2024)
por: Xu, Kevin, et al.
Publicado: (2024)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
por: Sun, Qiushi, et al.
Publicado: (2024)
por: Sun, Qiushi, et al.
Publicado: (2024)
SwissADT: An Audio Description Translation System for Swiss Languages
por: Fischer, Lukas, et al.
Publicado: (2024)
por: Fischer, Lukas, et al.
Publicado: (2024)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
por: Yin, Kayo, et al.
Publicado: (2024)
por: Yin, Kayo, et al.
Publicado: (2024)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
por: Lo, Leo Yu-Ho, et al.
Publicado: (2024)
por: Lo, Leo Yu-Ho, et al.
Publicado: (2024)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
por: Sun, Qiushi, et al.
Publicado: (2025)
por: Sun, Qiushi, et al.
Publicado: (2025)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
por: Wang, Yibo, et al.
Publicado: (2026)
por: Wang, Yibo, et al.
Publicado: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
por: Jing, Hongyi, et al.
Publicado: (2025)
por: Jing, Hongyi, et al.
Publicado: (2025)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
por: Wang, Haoming, et al.
Publicado: (2025)
por: Wang, Haoming, et al.
Publicado: (2025)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
por: Xie, Tianbao, et al.
Publicado: (2025)
por: Xie, Tianbao, et al.
Publicado: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
por: Ouyang, Mingyu, et al.
Publicado: (2026)
por: Ouyang, Mingyu, et al.
Publicado: (2026)
DesignPref: Capturing Personal Preferences in Visual Design Generation
por: Peng, Yi-Hao, et al.
Publicado: (2025)
por: Peng, Yi-Hao, et al.
Publicado: (2025)
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
por: Fan, Jingru, et al.
Publicado: (2025)
por: Fan, Jingru, et al.
Publicado: (2025)
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
por: Khalil, Mohammad Amer, et al.
Publicado: (2026)
por: Khalil, Mohammad Amer, et al.
Publicado: (2026)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
por: Cheng, Kanzhi, et al.
Publicado: (2026)
por: Cheng, Kanzhi, et al.
Publicado: (2026)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
por: Qin, Yujia, et al.
Publicado: (2025)
por: Qin, Yujia, et al.
Publicado: (2025)
History-Aware Reasoning for GUI Agents
por: Wang, Ziwei, et al.
Publicado: (2025)
por: Wang, Ziwei, et al.
Publicado: (2025)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
por: Chen, Xuetian, et al.
Publicado: (2025)
por: Chen, Xuetian, et al.
Publicado: (2025)
Forest-Chat: Adapting Vision-Language Agents for Interactive Forest Change Analysis
por: Brock, James, et al.
Publicado: (2026)
por: Brock, James, et al.
Publicado: (2026)
A Survey on (M)LLM-Based GUI Agents
por: Tang, Fei, et al.
Publicado: (2025)
por: Tang, Fei, et al.
Publicado: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
por: Wu, Zongru, et al.
Publicado: (2025)
por: Wu, Zongru, et al.
Publicado: (2025)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
por: Mukhopadhyay, Srija, et al.
Publicado: (2024)
por: Mukhopadhyay, Srija, et al.
Publicado: (2024)
Ejemplares similares
-
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
por: Zhao, Henry Hengyuan, et al.
Publicado: (2025) -
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
por: Sun, Qiushi, et al.
Publicado: (2025) -
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
por: Kapoor, Raghav, et al.
Publicado: (2024) -
AIN: The Arabic INclusive Large Multimodal Model
por: Heakl, Ahmed, et al.
Publicado: (2025) -
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
por: Lin, Ronghao, et al.
Publicado: (2025)