Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Lin, Xiu, Yanming, Gorlatova, Maria |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demonstrating Visual Information Manipulation Attacks in Augmented Reality: A Hands-On Miniature City-Based Setup
by: Xiu, Yanming, et al.
Published: (2025)
by: Xiu, Yanming, et al.
Published: (2025)
Do Vision Language Models Understand Human Engagement in Games?
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
by: Mukhopadhyay, Srija, et al.
Published: (2024)
by: Mukhopadhyay, Srija, et al.
Published: (2024)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)
by: Natalie, Rosiana, et al.
Published: (2025)
Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality
by: Xiu, Yanming, et al.
Published: (2025)
by: Xiu, Yanming, et al.
Published: (2025)
ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality
by: Xiu, Yanming, et al.
Published: (2025)
by: Xiu, Yanming, et al.
Published: (2025)
Refusal as Silence: Gendered Disparities in Vision-Language Model Responses
by: Luo, Sha, et al.
Published: (2024)
by: Luo, Sha, et al.
Published: (2024)
SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
by: Lin, Bo, et al.
Published: (2024)
by: Lin, Bo, et al.
Published: (2024)
Words into World: A Task-Adaptive Agent for Language-Guided Spatial Retrieval in AR
by: Guo, Lixing, et al.
Published: (2025)
by: Guo, Lixing, et al.
Published: (2025)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
GUICourse: From General Vision Language Models to Versatile GUI Agents
by: Chen, Wentong, et al.
Published: (2024)
by: Chen, Wentong, et al.
Published: (2024)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
by: Zhu, Ruijie, et al.
Published: (2025)
by: Zhu, Ruijie, et al.
Published: (2025)
Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
by: Bossen, Tonko E. W., et al.
Published: (2025)
by: Bossen, Tonko E. W., et al.
Published: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
by: Huang, Jinbin, et al.
Published: (2023)
by: Huang, Jinbin, et al.
Published: (2023)
Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
by: Jha, Saurav, et al.
Published: (2025)
by: Jha, Saurav, et al.
Published: (2025)
UI-UG: A Unified MLLM for UI Understanding and Generation
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024)
by: Niu, Runliang, et al.
Published: (2024)
Trust in Vision-Language Models: Insights from a Participatory User Workshop
by: Chiatti, Agnese, et al.
Published: (2025)
by: Chiatti, Agnese, et al.
Published: (2025)
Scene-Aware Urban Design: A Human-AI Recommendation Framework Using Co-Occurrence Embeddings and Vision-Language Models
by: Gallardo, Rodrigo, et al.
Published: (2025)
by: Gallardo, Rodrigo, et al.
Published: (2025)
Towards Context-aware Support for Color Vision Deficiency: An Approach Integrating LLM and AR
by: Morita, Shogo, et al.
Published: (2024)
by: Morita, Shogo, et al.
Published: (2024)
Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach
by: Xiu, Yanming, et al.
Published: (2025)
by: Xiu, Yanming, et al.
Published: (2025)
Constructive Apraxia: An Unexpected Limit of Instructible Vision-Language Models and Analog for Human Cognitive Disorders
by: Noever, David, et al.
Published: (2024)
by: Noever, David, et al.
Published: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Vision Language Models as Values Detectors
by: Abbo, Giulio Antonio, et al.
Published: (2025)
by: Abbo, Giulio Antonio, et al.
Published: (2025)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation
by: Kim, Namhee, et al.
Published: (2025)
by: Kim, Namhee, et al.
Published: (2025)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
by: Garg, Kapil, et al.
Published: (2025)
by: Garg, Kapil, et al.
Published: (2025)
Yume: An Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025)
by: Hu, Siyuan, et al.
Published: (2025)
Acoustic Field Video for Multimodal Scene Understanding
by: Kim, Daehwa, et al.
Published: (2026)
by: Kim, Daehwa, et al.
Published: (2026)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
by: Kondic, Jovana, et al.
Published: (2025)
by: Kondic, Jovana, et al.
Published: (2025)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
by: Huang, Yifei, et al.
Published: (2025)
by: Huang, Yifei, et al.
Published: (2025)
VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis
by: Yan, Xinyuan, et al.
Published: (2025)
by: Yan, Xinyuan, et al.
Published: (2025)
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
by: Han, Kyungtae, et al.
Published: (2025)
by: Han, Kyungtae, et al.
Published: (2025)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
by: Yang, Saelyne, et al.
Published: (2026)
by: Yang, Saelyne, et al.
Published: (2026)
When, Where, and What? A Novel Benchmark for Accident Anticipation and Localization with Large Language Models
by: Liao, Haicheng, et al.
Published: (2024)
by: Liao, Haicheng, et al.
Published: (2024)
User Prompting Strategies and Prompt Enhancement Methods for Open-Set Object Detection in XR Environments
by: Lin, Junfeng, et al.
Published: (2026)
by: Lin, Junfeng, et al.
Published: (2026)
Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality
by: Xiu, Yanming, et al.
Published: (2026)
by: Xiu, Yanming, et al.
Published: (2026)
Training a Vision Language Model as Smartphone Assistant
by: Dorka, Nicolai, et al.
Published: (2024)
by: Dorka, Nicolai, et al.
Published: (2024)
Similar Items
-
Demonstrating Visual Information Manipulation Attacks in Augmented Reality: A Hands-On Miniature City-Based Setup
by: Xiu, Yanming, et al.
Published: (2025) -
Do Vision Language Models Understand Human Engagement in Games?
by: Wang, Ziyi, et al.
Published: (2026) -
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
by: Mukhopadhyay, Srija, et al.
Published: (2024) -
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025) -
Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality
by: Xiu, Yanming, et al.
Published: (2025)