LingoQA: Visual Question Answering for Autonomous Driving
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913518824980480 |
|---|---|
| author | Marcu, Ana-Maria Chen, Long Hünermann, Jan Karnsund, Alice Hanotte, Benoit Chidananda, Prajwal Nair, Saurabh Badrinarayanan, Vijay Kendall, Alex Shotton, Jamie Arani, Elahe Sinavski, Oleg |
| author_facet | Marcu, Ana-Maria Chen, Long Hünermann, Jan Karnsund, Alice Hanotte, Benoit Chidananda, Prajwal Nair, Saurabh Badrinarayanan, Vijay Kendall, Alex Shotton, Jamie Arani, Elahe Sinavski, Oleg |
| contents | We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capabilities, with GPT-4V responding truthfully to 59.6% of the questions compared to 96.6% for humans. For evaluation, we propose a truthfulness classifier, called Lingo-Judge, that achieves a 0.95 Spearman correlation coefficient to human evaluations, surpassing existing techniques like METEOR, BLEU, CIDEr, and GPT-4. We establish a baseline vision-language model and run extensive ablation studies to understand its performance. We release our dataset and benchmark as an evaluation platform for vision-language models in autonomous driving. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_14115 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | LingoQA: Visual Question Answering for Autonomous Driving Marcu, Ana-Maria Chen, Long Hünermann, Jan Karnsund, Alice Hanotte, Benoit Chidananda, Prajwal Nair, Saurabh Badrinarayanan, Vijay Kendall, Alex Shotton, Jamie Arani, Elahe Sinavski, Oleg Robotics Artificial Intelligence Computer Vision and Pattern Recognition We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capabilities, with GPT-4V responding truthfully to 59.6% of the questions compared to 96.6% for humans. For evaluation, we propose a truthfulness classifier, called Lingo-Judge, that achieves a 0.95 Spearman correlation coefficient to human evaluations, surpassing existing techniques like METEOR, BLEU, CIDEr, and GPT-4. We establish a baseline vision-language model and run extensive ablation studies to understand its performance. We release our dataset and benchmark as an evaluation platform for vision-language models in autonomous driving. |
| title | LingoQA: Visual Question Answering for Autonomous Driving |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2312.14115 |