Gespeichert in:
| Hauptverfasser: | Pal, Ankit, Sankarasubbu, Malaikannan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2402.07023 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Review on Large Language Models for Visual Analytics
von: Agarwal, Navya Sonal, et al.
Veröffentlicht: (2025)
von: Agarwal, Navya Sonal, et al.
Veröffentlicht: (2025)
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
von: Mazumdar, Amrita, et al.
Veröffentlicht: (2026)
von: Mazumdar, Amrita, et al.
Veröffentlicht: (2026)
SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model
von: Cazzaniga, Luca
Veröffentlicht: (2026)
von: Cazzaniga, Luca
Veröffentlicht: (2026)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding
von: Pal, Ankit, et al.
Veröffentlicht: (2025)
von: Pal, Ankit, et al.
Veröffentlicht: (2025)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
von: You, Keen, et al.
Veröffentlicht: (2024)
von: You, Keen, et al.
Veröffentlicht: (2024)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
How Can Large Language Models Enable Better Socially Assistive Human-Robot Interaction: A Brief Survey
von: Shi, Zhonghao, et al.
Veröffentlicht: (2024)
von: Shi, Zhonghao, et al.
Veröffentlicht: (2024)
ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
von: Xia, Ding, et al.
Veröffentlicht: (2025)
von: Xia, Ding, et al.
Veröffentlicht: (2025)
Can Large Language Models Capture Video Game Engagement?
von: Melhart, David, et al.
Veröffentlicht: (2025)
von: Melhart, David, et al.
Veröffentlicht: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications
von: Asseri, Bushra, et al.
Veröffentlicht: (2025)
von: Asseri, Bushra, et al.
Veröffentlicht: (2025)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
Measuring Agreeableness Bias in Multimodal Models
von: Lim, Jaehyuk, et al.
Veröffentlicht: (2024)
von: Lim, Jaehyuk, et al.
Veröffentlicht: (2024)
AIN: The Arabic INclusive Large Multimodal Model
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
von: Betala, Siddharth, et al.
Veröffentlicht: (2025)
von: Betala, Siddharth, et al.
Veröffentlicht: (2025)
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
von: Lopez-Cardona, Angela, et al.
Veröffentlicht: (2024)
von: Lopez-Cardona, Angela, et al.
Veröffentlicht: (2024)
ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
von: Hsiao, Yu-Chung, et al.
Veröffentlicht: (2022)
von: Hsiao, Yu-Chung, et al.
Veröffentlicht: (2022)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
Learning Multimodal Cues of Children's Uncertainty
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
UIClip: A Data-driven Model for Assessing User Interface Design
von: Wu, Jason, et al.
Veröffentlicht: (2024)
von: Wu, Jason, et al.
Veröffentlicht: (2024)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
von: Yang, Kaichun, et al.
Veröffentlicht: (2025)
von: Yang, Kaichun, et al.
Veröffentlicht: (2025)
Voting-based Multimodal Automatic Deception Detection
von: Touma, Lana, et al.
Veröffentlicht: (2023)
von: Touma, Lana, et al.
Veröffentlicht: (2023)
Detoxifying Large Language Models via Knowledge Editing
von: Wang, Mengru, et al.
Veröffentlicht: (2024)
von: Wang, Mengru, et al.
Veröffentlicht: (2024)
Knowledge Mechanisms in Large Language Models: A Survey and Perspective
von: Wang, Mengru, et al.
Veröffentlicht: (2024)
von: Wang, Mengru, et al.
Veröffentlicht: (2024)
ReLearn: Unlearning via Learning for Large Language Models
von: Xu, Haoming, et al.
Veröffentlicht: (2025)
von: Xu, Haoming, et al.
Veröffentlicht: (2025)
A Comprehensive Study of Knowledge Editing for Large Language Models
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
GUICourse: From General Vision Language Models to Versatile GUI Agents
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
InstructEdit: Instruction-based Knowledge Editing for Large Language Models
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
Position and Rotation Invariant Sign Language Recognition from 3D Kinect Data with Recurrent Neural Networks
von: Roy, Prasun, et al.
Veröffentlicht: (2020)
von: Roy, Prasun, et al.
Veröffentlicht: (2020)
Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception
von: da Silva, Neemias, et al.
Veröffentlicht: (2026)
von: da Silva, Neemias, et al.
Veröffentlicht: (2026)
Semantic and Expressive Variation in Image Captions Across Languages
von: Ye, Andre, et al.
Veröffentlicht: (2023)
von: Ye, Andre, et al.
Veröffentlicht: (2023)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
MAP: Evaluation and Multi-Agent Enhancement of Large Language Models for Inpatient Pathways
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
von: Gao, Nan, et al.
Veröffentlicht: (2023)
von: Gao, Nan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A Review on Large Language Models for Visual Analytics
von: Agarwal, Navya Sonal, et al.
Veröffentlicht: (2025) -
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
von: Li, Xiao, et al.
Veröffentlicht: (2025) -
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
von: Mazumdar, Amrita, et al.
Veröffentlicht: (2026) -
SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model
von: Cazzaniga, Luca
Veröffentlicht: (2026) -
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)