LingoQA: Visual Question Answering for Autonomous Driving

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Marcu, Ana-Maria, Chen, Long, Hünermann, Jan, Karnsund, Alice, Hanotte, Benoit, Chidananda, Prajwal, Nair, Saurabh, Badrinarayanan, Vijay, Kendall, Alex, Shotton, Jamie, Arani, Elahe, Sinavski, Oleg
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913518824980480
author Marcu, Ana-Maria
Chen, Long
Hünermann, Jan
Karnsund, Alice
Hanotte, Benoit
Chidananda, Prajwal
Nair, Saurabh
Badrinarayanan, Vijay
Kendall, Alex
Shotton, Jamie
Arani, Elahe
Sinavski, Oleg
author_facet Marcu, Ana-Maria
Chen, Long
Hünermann, Jan
Karnsund, Alice
Hanotte, Benoit
Chidananda, Prajwal
Nair, Saurabh
Badrinarayanan, Vijay
Kendall, Alex
Shotton, Jamie
Arani, Elahe
Sinavski, Oleg
contents We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capabilities, with GPT-4V responding truthfully to 59.6% of the questions compared to 96.6% for humans. For evaluation, we propose a truthfulness classifier, called Lingo-Judge, that achieves a 0.95 Spearman correlation coefficient to human evaluations, surpassing existing techniques like METEOR, BLEU, CIDEr, and GPT-4. We establish a baseline vision-language model and run extensive ablation studies to understand its performance. We release our dataset and benchmark as an evaluation platform for vision-language models in autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14115
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LingoQA: Visual Question Answering for Autonomous Driving
Marcu, Ana-Maria
Chen, Long
Hünermann, Jan
Karnsund, Alice
Hanotte, Benoit
Chidananda, Prajwal
Nair, Saurabh
Badrinarayanan, Vijay
Kendall, Alex
Shotton, Jamie
Arani, Elahe
Sinavski, Oleg
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capabilities, with GPT-4V responding truthfully to 59.6% of the questions compared to 96.6% for humans. For evaluation, we propose a truthfulness classifier, called Lingo-Judge, that achieves a 0.95 Spearman correlation coefficient to human evaluations, surpassing existing techniques like METEOR, BLEU, CIDEr, and GPT-4. We establish a baseline vision-language model and run extensive ablation studies to understand its performance. We release our dataset and benchmark as an evaluation platform for vision-language models in autonomous driving.
title LingoQA: Visual Question Answering for Autonomous Driving
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.14115