VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
Fuente:
arXiv
Saved in:
| Main Authors: | Taesiri, Mohammad Reza, Ghildyal, Abhijay, Zadtootaghaj, Saman, Barman, Nabajeet, Bezemer, Cor-Paul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Foundation Models Boost Low-Level Perceptual Similarity Metrics
by: Ghildyal, Abhijay, et al.
Published: (2024)
by: Ghildyal, Abhijay, et al.
Published: (2024)
Quality Prediction of AI Generated Images and Videos: Emerging Trends and Opportunities
by: Ghildyal, Abhijay, et al.
Published: (2024)
by: Ghildyal, Abhijay, et al.
Published: (2024)
MSLIQA: Enhancing Learning Representations for Image Quality Assessment through Multi-Scale Learning
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
LAR-IQA: A Lightweight, Accurate, and Robust No-Reference Image Quality Assessment Model
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
Non-Aligned Reference Image Quality Assessment for Novel View Synthesis
by: Ghildyal, Abhijay, et al.
Published: (2025)
by: Ghildyal, Abhijay, et al.
Published: (2025)
VideoGameBunny: Towards vision assistants for video games
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
RESP: Reference-guided Sequential Prompting for Visual Glitch Detection in Video Games
by: Yu, Yakun, et al.
Published: (2026)
by: Yu, Yakun, et al.
Published: (2026)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
by: Yu, Yakun, et al.
Published: (2026)
by: Yu, Yakun, et al.
Published: (2026)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
TRIQA: Image Quality Assessment by Contrastive Pretraining on Ordered Distortion Triplets
by: Sureddi, Rajesh, et al.
Published: (2025)
by: Sureddi, Rajesh, et al.
Published: (2025)
AIM 2024 Challenge on UHD Blind Photo Quality Assessment
by: Hosu, Vlad, et al.
Published: (2024)
by: Hosu, Vlad, et al.
Published: (2024)
WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
by: Ghildyal, Abhijay, et al.
Published: (2025)
by: Ghildyal, Abhijay, et al.
Published: (2025)
How Far Can VLMs Go for Visual Bug Detection? Studying 19,738 Keyframes from 41 Hours of Gameplay Videos
by: Lu, Wentao, et al.
Published: (2026)
by: Lu, Wentao, et al.
Published: (2026)
VideoGameBench: Can Vision-Language Models complete popular video games?
by: Zhang, Alex L., et al.
Published: (2025)
by: Zhang, Alex L., et al.
Published: (2025)
Vision Language Models are Biased
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
by: Zhang, Kuan, et al.
Published: (2026)
by: Zhang, Kuan, et al.
Published: (2026)
Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
by: Lu, Wentao, et al.
Published: (2025)
by: Lu, Wentao, et al.
Published: (2025)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
Exploring the Capabilities of Vision-Language Models to Detect Visual Bugs in HTML5 <canvas> Applications
by: Macklon, Finlay, et al.
Published: (2025)
by: Macklon, Finlay, et al.
Published: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
by: Li, Chenglin, et al.
Published: (2024)
by: Li, Chenglin, et al.
Published: (2024)
GameFactory: Creating New Games with Generative Interactive Videos
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Study of Subjective and Objective Quality Assessment of Mobile Cloud Gaming Videos
by: Saha, Avinab, et al.
Published: (2023)
by: Saha, Avinab, et al.
Published: (2023)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
AIS 2024 Challenge on Video Quality Assessment of User-Generated Content: Methods and Results
by: Conde, Marcos V., et al.
Published: (2024)
by: Conde, Marcos V., et al.
Published: (2024)
GameScope: A Multi-Attribute, Multi-Codec Benchmark Dataset for Gaming Video Quality Assessment
by: Sureddi, Rajesh, et al.
Published: (2026)
by: Sureddi, Rajesh, et al.
Published: (2026)
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
GameGen-X: Interactive Open-world Game Video Generation
by: Che, Haoxuan, et al.
Published: (2024)
by: Che, Haoxuan, et al.
Published: (2024)
The Weighting Game: Evaluating Quality of Explainability Methods
by: Raatikainen, Lassi, et al.
Published: (2022)
by: Raatikainen, Lassi, et al.
Published: (2022)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Can Large Language Models Capture Video Game Engagement?
by: Melhart, David, et al.
Published: (2025)
by: Melhart, David, et al.
Published: (2025)
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
by: Ghoddoosian, Reza, et al.
Published: (2024)
by: Ghoddoosian, Reza, et al.
Published: (2024)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
by: Zhang, Zhihong, et al.
Published: (2025)
by: Zhang, Zhihong, et al.
Published: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026)
by: Collins, Brandon, et al.
Published: (2026)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
Can Vision-Language Models Solve the Shell Game?
by: Liu, Tiedong, et al.
Published: (2026)
by: Liu, Tiedong, et al.
Published: (2026)
Similar Items
-
Foundation Models Boost Low-Level Perceptual Similarity Metrics
by: Ghildyal, Abhijay, et al.
Published: (2024) -
Quality Prediction of AI Generated Images and Videos: Emerging Trends and Opportunities
by: Ghildyal, Abhijay, et al.
Published: (2024) -
MSLIQA: Enhancing Learning Representations for Image Quality Assessment through Multi-Scale Learning
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024) -
LAR-IQA: A Lightweight, Accurate, and Robust No-Reference Image Quality Assessment Model
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024) -
Non-Aligned Reference Image Quality Assessment for Novel View Synthesis
by: Ghildyal, Abhijay, et al.
Published: (2025)