ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Baechler, Gilles, Sunkara, Srinivas, Wang, Maria, Zubach, Fedir, Mansoor, Hassan, Etter, Vincent, Cărbune, Victor, Lin, Jason, Chen, Jindong, Sharma, Abhanshu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
by: Hsiao, Yu-Chung, et al.
Published: (2022)
by: Hsiao, Yu-Chung, et al.
Published: (2022)
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
by: Carbune, Victor, et al.
Published: (2024)
by: Carbune, Victor, et al.
Published: (2024)
WebQuest: A Benchmark for Multimodal QA on Web Page Sequences
by: Wang, Maria, et al.
Published: (2024)
by: Wang, Maria, et al.
Published: (2024)
VQA Training Sets are Self-play Environments for Generating Few-shot Pools
by: Misiunas, Tautvydas, et al.
Published: (2024)
by: Misiunas, Tautvydas, et al.
Published: (2024)
LLMs cannot find reasoning errors, but can correct them given the error location
by: Tyen, Gladys, et al.
Published: (2023)
by: Tyen, Gladys, et al.
Published: (2023)
UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments
by: Xiang, Jiannan, et al.
Published: (2025)
by: Xiang, Jiannan, et al.
Published: (2025)
Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
by: Hazimeh, Adam, et al.
Published: (2025)
by: Hazimeh, Adam, et al.
Published: (2025)
Absolute Komplexität in der Nominalflexion
by: Baechler, Raffaela
Published: (2018)
by: Baechler, Raffaela
Published: (2018)
Nicht-phonologisch konditionierter Wandel in der Kasusmorphologie isolierter germanischer Varietäten. Höchstalemannisch (Visperterminen) und Älvdalisch
by: Raffaela Baechler
Published: (2019)
by: Raffaela Baechler
Published: (2019)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
by: Lee, Harrison, et al.
Published: (2023)
by: Lee, Harrison, et al.
Published: (2023)
The effect of vinboron on the expression processes of apoptosis in gastric mucosa with ibuprofen–induced gastropathy in rats
by: Fedir Hladkykh
Published: (2016)
by: Fedir Hladkykh
Published: (2016)
Vikings in the East
by: Androshchuk, Fedir
Published: (2025)
by: Androshchuk, Fedir
Published: (2025)
AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding
by: Khan, Safwat Ali, et al.
Published: (2024)
by: Khan, Safwat Ali, et al.
Published: (2024)
The Language of Infographics: Toward Understanding Conceptual Metaphor Use in Scientific Storytelling
by: Pokojná, Hana, et al.
Published: (2024)
by: Pokojná, Hana, et al.
Published: (2024)
UI-UG: A Unified MLLM for UI Understanding and Generation
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
by: Wu, Qinzhuo, et al.
Published: (2024)
by: Wu, Qinzhuo, et al.
Published: (2024)
From Dead Pixels to Editable Slides: Infographic Reconstruction into Native Google Slides via Vision-Language Region Understanding
by: Gonzalez, Leonardo
Published: (2026)
by: Gonzalez, Leonardo
Published: (2026)
An Evaluation of Immersive Infographics for News Reporting: Quantifying the Effect of Mobile AR Concrete Scales Infographics on Volume Understanding
by: Giambastiani, Mariane, et al.
Published: (2024)
by: Giambastiani, Mariane, et al.
Published: (2024)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
by: Yang, Jiaxi, et al.
Published: (2025)
by: Yang, Jiaxi, et al.
Published: (2025)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
by: Sunkara, Krishna Chaitanya
Published: (2026)
by: Sunkara, Krishna Chaitanya
Published: (2026)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
Fig. 3 in A report on palaeontological excavations and sampling in mudrocks: some guidelines
by: Etter, Walter
Published: (2024)
by: Etter, Walter
Published: (2024)
The brain as a blueprint: a survey of brain-inspired approaches to learning in artificial intelligence
by: Etter, Guillaume
Published: (2025)
by: Etter, Guillaume
Published: (2025)
Motorola Infographics
by: Motorola
Published: (2026)
by: Motorola
Published: (2026)
February Infographic
Published: (2025)
Published: (2025)
September Infographic
Published: (2025)
Published: (2025)
May Infographic
Published: (2025)
Published: (2025)
March Infographic
Published: (2026)
Published: (2026)
April Infographic
Published: (2024)
Published: (2024)
March Infographic
Published: (2024)
Published: (2024)
April Infographic
Published: (2025)
Published: (2025)
May Infographic
Published: (2026)
Published: (2026)
Similar Items
-
ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
by: Hsiao, Yu-Chung, et al.
Published: (2022) -
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
by: Carbune, Victor, et al.
Published: (2024) -
WebQuest: A Benchmark for Multimodal QA on Web Page Sequences
by: Wang, Maria, et al.
Published: (2024) -
VQA Training Sets are Self-play Environments for Generating Few-shot Pools
by: Misiunas, Tautvydas, et al.
Published: (2024) -
LLMs cannot find reasoning errors, but can correct them given the error location
by: Tyen, Gladys, et al.
Published: (2023)