Salvato in:
| Autori principali: | Lim, Jaehyuk, Lee, Bruce W. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2408.09111 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
di: Zhao, Henry Hengyuan, et al.
Pubblicazione: (2025)
di: Zhao, Henry Hengyuan, et al.
Pubblicazione: (2025)
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
di: Kapoor, Raghav, et al.
Pubblicazione: (2024)
di: Kapoor, Raghav, et al.
Pubblicazione: (2024)
AIN: The Arabic INclusive Large Multimodal Model
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
di: Tsangko, Iosif, et al.
Pubblicazione: (2025)
di: Tsangko, Iosif, et al.
Pubblicazione: (2025)
Can Large Language Models Capture Video Game Engagement?
di: Melhart, David, et al.
Pubblicazione: (2025)
di: Melhart, David, et al.
Pubblicazione: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
di: Chen, Wentong, et al.
Pubblicazione: (2024)
di: Chen, Wentong, et al.
Pubblicazione: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
di: Chen, Baiyu, et al.
Pubblicazione: (2026)
di: Chen, Baiyu, et al.
Pubblicazione: (2026)
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
di: Huang, Ting-Hao 'Kenneth', et al.
Pubblicazione: (2025)
di: Huang, Ting-Hao 'Kenneth', et al.
Pubblicazione: (2025)
Code2World: A GUI World Model via Renderable Code Generation
di: Zheng, Yuhao, et al.
Pubblicazione: (2026)
di: Zheng, Yuhao, et al.
Pubblicazione: (2026)
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
di: Lopez-Cardona, Angela, et al.
Pubblicazione: (2024)
di: Lopez-Cardona, Angela, et al.
Pubblicazione: (2024)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
di: Tsangko, Iosif, et al.
Pubblicazione: (2026)
di: Tsangko, Iosif, et al.
Pubblicazione: (2026)
Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations
di: Pal, Ankit, et al.
Pubblicazione: (2024)
di: Pal, Ankit, et al.
Pubblicazione: (2024)
EEG-based Multimodal Representation Learning for Emotion Recognition
di: Yin, Kang, et al.
Pubblicazione: (2024)
di: Yin, Kang, et al.
Pubblicazione: (2024)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
di: Xu, Kevin, et al.
Pubblicazione: (2024)
di: Xu, Kevin, et al.
Pubblicazione: (2024)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
di: Sun, Qiushi, et al.
Pubblicazione: (2024)
di: Sun, Qiushi, et al.
Pubblicazione: (2024)
SwissADT: An Audio Description Translation System for Swiss Languages
di: Fischer, Lukas, et al.
Pubblicazione: (2024)
di: Fischer, Lukas, et al.
Pubblicazione: (2024)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
di: Yin, Kayo, et al.
Pubblicazione: (2024)
di: Yin, Kayo, et al.
Pubblicazione: (2024)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
di: Lo, Leo Yu-Ho, et al.
Pubblicazione: (2024)
di: Lo, Leo Yu-Ho, et al.
Pubblicazione: (2024)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
di: Wang, Yibo, et al.
Pubblicazione: (2026)
di: Wang, Yibo, et al.
Pubblicazione: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
di: Jing, Hongyi, et al.
Pubblicazione: (2025)
di: Jing, Hongyi, et al.
Pubblicazione: (2025)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
di: Wang, Haoming, et al.
Pubblicazione: (2025)
di: Wang, Haoming, et al.
Pubblicazione: (2025)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
di: Xie, Tianbao, et al.
Pubblicazione: (2025)
di: Xie, Tianbao, et al.
Pubblicazione: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
di: Ouyang, Mingyu, et al.
Pubblicazione: (2026)
di: Ouyang, Mingyu, et al.
Pubblicazione: (2026)
DesignPref: Capturing Personal Preferences in Visual Design Generation
di: Peng, Yi-Hao, et al.
Pubblicazione: (2025)
di: Peng, Yi-Hao, et al.
Pubblicazione: (2025)
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
di: Fan, Jingru, et al.
Pubblicazione: (2025)
di: Fan, Jingru, et al.
Pubblicazione: (2025)
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
di: Khalil, Mohammad Amer, et al.
Pubblicazione: (2026)
di: Khalil, Mohammad Amer, et al.
Pubblicazione: (2026)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
di: Cheng, Kanzhi, et al.
Pubblicazione: (2026)
di: Cheng, Kanzhi, et al.
Pubblicazione: (2026)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
di: Qin, Yujia, et al.
Pubblicazione: (2025)
di: Qin, Yujia, et al.
Pubblicazione: (2025)
History-Aware Reasoning for GUI Agents
di: Wang, Ziwei, et al.
Pubblicazione: (2025)
di: Wang, Ziwei, et al.
Pubblicazione: (2025)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
di: Chen, Xuetian, et al.
Pubblicazione: (2025)
di: Chen, Xuetian, et al.
Pubblicazione: (2025)
Forest-Chat: Adapting Vision-Language Agents for Interactive Forest Change Analysis
di: Brock, James, et al.
Pubblicazione: (2026)
di: Brock, James, et al.
Pubblicazione: (2026)
A Survey on (M)LLM-Based GUI Agents
di: Tang, Fei, et al.
Pubblicazione: (2025)
di: Tang, Fei, et al.
Pubblicazione: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
di: Wu, Zongru, et al.
Pubblicazione: (2025)
di: Wu, Zongru, et al.
Pubblicazione: (2025)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2024)
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2024)
Documenti analoghi
-
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
di: Zhao, Henry Hengyuan, et al.
Pubblicazione: (2025) -
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
di: Sun, Qiushi, et al.
Pubblicazione: (2025) -
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
di: Kapoor, Raghav, et al.
Pubblicazione: (2024) -
AIN: The Arabic INclusive Large Multimodal Model
di: Heakl, Ahmed, et al.
Pubblicazione: (2025) -
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
di: Lin, Ronghao, et al.
Pubblicazione: (2025)