Salvato in:
| Autori principali: | Luo, Sha, Kim, Sang Jung, Duan, Zening, Chen, Kaiping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2406.08222 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
Vision-Language Models Suppress Female Representations Under Ambiguous Input
di: Marin-Llobet, Arnau, et al.
Pubblicazione: (2026)
di: Marin-Llobet, Arnau, et al.
Pubblicazione: (2026)
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
di: Duan, Lin, et al.
Pubblicazione: (2025)
di: Duan, Lin, et al.
Pubblicazione: (2025)
BLK-Assist: A Methodological Framework for Artist-Led Co-Creation with Generative AI Models
di: Grimes, Daniel, et al.
Pubblicazione: (2026)
di: Grimes, Daniel, et al.
Pubblicazione: (2026)
Using Salient Object Detection to Identify Manipulative Cookie Banners that Circumvent GDPR
di: Grossman, Riley, et al.
Pubblicazione: (2025)
di: Grossman, Riley, et al.
Pubblicazione: (2025)
Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
di: Yang, Sicheng, et al.
Pubblicazione: (2026)
di: Yang, Sicheng, et al.
Pubblicazione: (2026)
Joining Forces for Pathology Diagnostics with AI Assistance: The EMPAIA Initiative
di: Zerbe, Norman, et al.
Pubblicazione: (2023)
di: Zerbe, Norman, et al.
Pubblicazione: (2023)
AI-based Multimodal Biometrics for Detecting Smartphone Distractions: Application to Online Learning
di: Becerra, Alvaro, et al.
Pubblicazione: (2025)
di: Becerra, Alvaro, et al.
Pubblicazione: (2025)
Gaze patterns predict preference and confidence in pairwise AI image evaluation
di: Papadopoulos, Nikolas, et al.
Pubblicazione: (2026)
di: Papadopoulos, Nikolas, et al.
Pubblicazione: (2026)
Negative Shanshui: Real-time Interactive Ink Painting Synthesis
di: Zhou, Aven-Le
Pubblicazione: (2025)
di: Zhou, Aven-Le
Pubblicazione: (2025)
AI-Based Facial Emotion Recognition Solutions for Education: A Study of Teacher-User and Other Categories
di: Ravenor, R. Yamamoto
Pubblicazione: (2023)
di: Ravenor, R. Yamamoto
Pubblicazione: (2023)
From Image Generation to Infrastructure Design: a Multi-agent Pipeline for Street Design Generation
di: Wang, Chenguang, et al.
Pubblicazione: (2025)
di: Wang, Chenguang, et al.
Pubblicazione: (2025)
Mask-up: Investigating Biases in Face Re-identification for Masked Faces
di: Jaiswal, Siddharth D, et al.
Pubblicazione: (2024)
di: Jaiswal, Siddharth D, et al.
Pubblicazione: (2024)
Handwritten Code Recognition for Pen-and-Paper CS Education
di: Islam, Md Sazzad, et al.
Pubblicazione: (2024)
di: Islam, Md Sazzad, et al.
Pubblicazione: (2024)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
di: Natalie, Rosiana, et al.
Pubblicazione: (2025)
di: Natalie, Rosiana, et al.
Pubblicazione: (2025)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
di: Duan, Ranjie, et al.
Pubblicazione: (2025)
di: Duan, Ranjie, et al.
Pubblicazione: (2025)
Do Vision Language Models Understand Human Engagement in Games?
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
di: Luo, Cheng, et al.
Pubblicazione: (2025)
di: Luo, Cheng, et al.
Pubblicazione: (2025)
ScreenAgent: A Vision Language Model-driven Computer Control Agent
di: Niu, Runliang, et al.
Pubblicazione: (2024)
di: Niu, Runliang, et al.
Pubblicazione: (2024)
Trust in Vision-Language Models: Insights from a Participatory User Workshop
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
di: Chen, Wentong, et al.
Pubblicazione: (2024)
di: Chen, Wentong, et al.
Pubblicazione: (2024)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
di: Hall, Melissa, et al.
Pubblicazione: (2024)
di: Hall, Melissa, et al.
Pubblicazione: (2024)
Improved Digital Therapy for Developmental Pediatrics Using Domain-Specific Artificial Intelligence: Machine Learning Study
di: Washington, Peter, et al.
Pubblicazione: (2020)
di: Washington, Peter, et al.
Pubblicazione: (2020)
Constructive Apraxia: An Unexpected Limit of Instructible Vision-Language Models and Analog for Human Cognitive Disorders
di: Noever, David, et al.
Pubblicazione: (2024)
di: Noever, David, et al.
Pubblicazione: (2024)
HarassGuard: Detecting Harassment Behaviors in Social Virtual Reality with Vision-Language Models
di: Lee, Junhee, et al.
Pubblicazione: (2026)
di: Lee, Junhee, et al.
Pubblicazione: (2026)
Generating Robot Constitutions & Benchmarks for Semantic Safety
di: Sermanet, Pierre, et al.
Pubblicazione: (2025)
di: Sermanet, Pierre, et al.
Pubblicazione: (2025)
Determining the Difficulties of Students With Dyslexia via Virtual Reality and Artificial Intelligence: An Exploratory Analysis
di: Yeguas-Bolívar, Enrique, et al.
Pubblicazione: (2024)
di: Yeguas-Bolívar, Enrique, et al.
Pubblicazione: (2024)
A Survey on Trustworthiness in Foundation Models for Medical Image Analysis
di: Shi, Congzhen, et al.
Pubblicazione: (2024)
di: Shi, Congzhen, et al.
Pubblicazione: (2024)
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation
di: Kim, Namhee, et al.
Pubblicazione: (2025)
di: Kim, Namhee, et al.
Pubblicazione: (2025)
Posture-Informed Muscular Force Learning for Robust Hand Pressure Estimation
di: Seo, Kyungjin, et al.
Pubblicazione: (2024)
di: Seo, Kyungjin, et al.
Pubblicazione: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
Vision Language Models as Values Detectors
di: Abbo, Giulio Antonio, et al.
Pubblicazione: (2025)
di: Abbo, Giulio Antonio, et al.
Pubblicazione: (2025)
Semantic and Expressive Variation in Image Captions Across Languages
di: Ye, Andre, et al.
Pubblicazione: (2023)
di: Ye, Andre, et al.
Pubblicazione: (2023)
SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
di: Lin, Bo, et al.
Pubblicazione: (2024)
di: Lin, Bo, et al.
Pubblicazione: (2024)
A Call to Arms: AI Should be Critical for Social Media Analysis of Conflict Zones
di: Abedin, Afia, et al.
Pubblicazione: (2023)
di: Abedin, Afia, et al.
Pubblicazione: (2023)
A Comparison of Human and Machine Learning Errors in Face Recognition
di: Estévez-Almenzar, Marina, et al.
Pubblicazione: (2025)
di: Estévez-Almenzar, Marina, et al.
Pubblicazione: (2025)
Classification of the lunar surface pattern by AI architectures: Does AI see a rabbit in the Moon?
di: Shoji, Daigo
Pubblicazione: (2023)
di: Shoji, Daigo
Pubblicazione: (2023)
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
di: Marquez-Carpintero, Luis, et al.
Pubblicazione: (2025)
di: Marquez-Carpintero, Luis, et al.
Pubblicazione: (2025)
Beyond Questionnaires: Video Analysis for Social Anxiety Detection
di: Sahu, Nilesh Kumar, et al.
Pubblicazione: (2024)
di: Sahu, Nilesh Kumar, et al.
Pubblicazione: (2024)
DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection
di: Ye, Jiaxin, et al.
Pubblicazione: (2024)
di: Ye, Jiaxin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects
di: Chiatti, Agnese, et al.
Pubblicazione: (2025) -
Vision-Language Models Suppress Female Representations Under Ambiguous Input
di: Marin-Llobet, Arnau, et al.
Pubblicazione: (2026) -
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
di: Duan, Lin, et al.
Pubblicazione: (2025) -
BLK-Assist: A Methodological Framework for Artist-Led Co-Creation with Generative AI Models
di: Grimes, Daniel, et al.
Pubblicazione: (2026) -
Using Salient Object Detection to Identify Manipulative Cookie Banners that Circumvent GDPR
di: Grossman, Riley, et al.
Pubblicazione: (2025)