How Far Can VLMs Go for Visual Bug Detection? Studying 19,738 Keyframes from 41 Hours of Gameplay Videos
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lu, Wentao, Senchenko, Alexander, Sayle, Alan, Hindle, Abram, Bezemer, Cor-Paul |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
par: Lu, Wentao, et autres
Publié: (2025)
par: Lu, Wentao, et autres
Publié: (2025)
Exploring the Capabilities of Vision-Language Models to Detect Visual Bugs in HTML5 <canvas> Applications
par: Macklon, Finlay, et autres
Publié: (2025)
par: Macklon, Finlay, et autres
Publié: (2025)
VideoGameBunny: Towards vision assistants for video games
par: Taesiri, Mohammad Reza, et autres
Publié: (2024)
par: Taesiri, Mohammad Reza, et autres
Publié: (2024)
Bridging the Language Gap: An Empirical Study of Bindings for Open Source Machine Learning Libraries Across Software Package Ecosystems
par: Li, Hao, et autres
Publié: (2022)
par: Li, Hao, et autres
Publié: (2022)
How do Agents Refactor: An Empirical Study
par: Ottenhof, Lukas, et autres
Publié: (2026)
par: Ottenhof, Lukas, et autres
Publié: (2026)
Patterns of Multi-Container Composition for Service Orchestration with Docker Compose
par: Eng, Kalvin, et autres
Publié: (2023)
par: Eng, Kalvin, et autres
Publié: (2023)
Software Engineering and Foundation Models: Insights from Industry Blogs Using a Jury of Foundation Models
par: Li, Hao, et autres
Publié: (2024)
par: Li, Hao, et autres
Publié: (2024)
Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
par: Li, Hao, et autres
Publié: (2024)
par: Li, Hao, et autres
Publié: (2024)
Assessing the Feasibility of Selective Instrumentation for Runtime Code Coverage in Large C++ Game Engines
par: Gauk, Ian, et autres
Publié: (2026)
par: Gauk, Ian, et autres
Publié: (2026)
Mining Type Constructs Using Patterns in AI-Generated Code
par: Lee, Imgyeong, et autres
Publié: (2026)
par: Lee, Imgyeong, et autres
Publié: (2026)
A Systematic Literature Review of Software Engineering Research on Jupyter Notebook
par: Siddik, Md Saeed, et autres
Publié: (2025)
par: Siddik, Md Saeed, et autres
Publié: (2025)
How Far Can We Go with Practical Function-Level Program Repair?
par: Xiang, Jiahong, et autres
Publié: (2024)
par: Xiang, Jiahong, et autres
Publié: (2024)
XBIDetective: Leveraging Vision Language Models for Identifying Cross-Browser Visual Inconsistencies
par: Grewal, Balreet, et autres
Publié: (2025)
par: Grewal, Balreet, et autres
Publié: (2025)
Duplicate Bug Report Detection: How Far Are We?
par: Zhang, Ting, et autres
Publié: (2022)
par: Zhang, Ting, et autres
Publié: (2022)
IRJIT: A Simple, Online, Information Retrieval Approach for Just-In-Time Software Defect Prediction
par: Sahar, Hareem, et autres
Publié: (2022)
par: Sahar, Hareem, et autres
Publié: (2022)
A Taxonomy of Testable HTML5 Canvas Issues
par: Macklon, Finlay, et autres
Publié: (2022)
par: Macklon, Finlay, et autres
Publié: (2022)
Keeping Deep Learning Models in Check: A History-Based Approach to Mitigate Overfitting
par: Li, Hao, et autres
Publié: (2024)
par: Li, Hao, et autres
Publié: (2024)
Early Detection of Performance Regressions by Bridging Local Performance Data and Architectural Models
par: Liao, Lizhi, et autres
Publié: (2024)
par: Liao, Lizhi, et autres
Publié: (2024)
How and Why Agents Can Identify Bug-Introducing Commits
par: Risse, Niklas, et autres
Publié: (2026)
par: Risse, Niklas, et autres
Publié: (2026)
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
par: Tsimpourlas, Foivos, et autres
Publié: (2024)
par: Tsimpourlas, Foivos, et autres
Publié: (2024)
Can AI Agents Generate Microservices? How Far are We?
par: Adnan, Bassam, et autres
Publié: (2026)
par: Adnan, Bassam, et autres
Publié: (2026)
Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?
par: Li, Shuqing, et autres
Publié: (2025)
par: Li, Shuqing, et autres
Publié: (2025)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
par: Yu, Yakun, et autres
Publié: (2026)
par: Yu, Yakun, et autres
Publié: (2026)
How Low Can You Go? The Data-Light SE Challenge
par: Ganguly, Kishan Kumar, et autres
Publié: (2025)
par: Ganguly, Kishan Kumar, et autres
Publié: (2025)
Exposing Go's Hidden Bugs: A Novel Concolic Framework
par: Gorna, Karolina, et autres
Publié: (2025)
par: Gorna, Karolina, et autres
Publié: (2025)
ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch
par: Si, Yuan, et autres
Publié: (2025)
par: Si, Yuan, et autres
Publié: (2025)
How is Testing Related to Single Statement Bugs?
par: Rahman, Habibur, et autres
Publié: (2024)
par: Rahman, Habibur, et autres
Publié: (2024)
TriGraph: A Probabilistic Subgraph-Based Model for Visual Code Completion in Pure Data (Package)
par: Islam, Anisha, et autres
Publié: (2024)
par: Islam, Anisha, et autres
Publié: (2024)
Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs
par: Hu, Haichuan, et autres
Publié: (2024)
par: Hu, Haichuan, et autres
Publié: (2024)
Large Language Models in Game Development: Implications for Gameplay, Playability, and Player Experience
par: Johnson, Keeryn, et autres
Publié: (2026)
par: Johnson, Keeryn, et autres
Publié: (2026)
Bug Whispering: Towards Audio Bug Reporting
par: Masserini, Elena, et autres
Publié: (2025)
par: Masserini, Elena, et autres
Publié: (2025)
Vulnerability-Affected Versions Identification: How Far Are We?
par: Chen, Xingchu, et autres
Publié: (2025)
par: Chen, Xingchu, et autres
Publié: (2025)
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
par: Acharya, Jagrit, et autres
Publié: (2025)
par: Acharya, Jagrit, et autres
Publié: (2025)
BugScope: Learn to Find Bugs Like Human
par: Guo, Jinyao, et autres
Publié: (2025)
par: Guo, Jinyao, et autres
Publié: (2025)
Towards Bug-Free Distributed Go Programs
par: Koo, Zhengqun
Publié: (2025)
par: Koo, Zhengqun
Publié: (2025)
Representation Learning for Stack Overflow Posts: How Far are We?
par: He, Junda, et autres
Publié: (2023)
par: He, Junda, et autres
Publié: (2023)
Automated Testing of Task-based Chatbots: How Far Are We?
par: Clerissi, Diego, et autres
Publié: (2026)
par: Clerissi, Diego, et autres
Publié: (2026)
Anticipating Bugs: Ticket-Level Bug Prediction and Temporal Proximity Effects
par: La Prova, Daniele, et autres
Publié: (2025)
par: La Prova, Daniele, et autres
Publié: (2025)
Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection
par: Rathnasuriya, Ravishka, et autres
Publié: (2026)
par: Rathnasuriya, Ravishka, et autres
Publié: (2026)
GitBug-Java: A Reproducible Benchmark of Recent Java Bugs
par: Silva, André, et autres
Publié: (2024)
par: Silva, André, et autres
Publié: (2024)
Documents similaires
-
Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
par: Lu, Wentao, et autres
Publié: (2025) -
Exploring the Capabilities of Vision-Language Models to Detect Visual Bugs in HTML5 <canvas> Applications
par: Macklon, Finlay, et autres
Publié: (2025) -
VideoGameBunny: Towards vision assistants for video games
par: Taesiri, Mohammad Reza, et autres
Publié: (2024) -
Bridging the Language Gap: An Empirical Study of Bindings for Open Source Machine Learning Libraries Across Software Package Ecosystems
par: Li, Hao, et autres
Publié: (2022) -
How do Agents Refactor: An Empirical Study
par: Ottenhof, Lukas, et autres
Publié: (2026)