Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Chih-Kai, Huang, Kuan-Po, Lee, Hung-yi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910605036748800
author Yang, Chih-Kai
Huang, Kuan-Po
Lee, Hung-yi
author_facet Yang, Chih-Kai
Huang, Kuan-Po
Lee, Hung-yi
contents This research explores how the information of prompts interacts with the high-performing speech recognition model, Whisper. We compare its performances when prompted by prompts with correct information and those corrupted with incorrect information. Our results unexpectedly show that Whisper may not understand the textual prompts in a human-expected way. Additionally, we find that performance improvement is not guaranteed even with stronger adherence to the topic information in textual prompts. It is also noted that English prompts generally outperform Mandarin ones on datasets of both languages, likely due to differences in training data distributions for these languages despite the mismatch with pre-training scenarios. Conversely, we discover that Whisper exhibits awareness of misleading information in language tokens by ignoring incorrect language tokens and focusing on the correct ones. In sum, We raise insightful questions about Whisper's prompt understanding and reveal its counter-intuitive behaviors. We encourage further studies.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05806
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
Yang, Chih-Kai
Huang, Kuan-Po
Lee, Hung-yi
Computation and Language
Sound
Audio and Speech Processing
This research explores how the information of prompts interacts with the high-performing speech recognition model, Whisper. We compare its performances when prompted by prompts with correct information and those corrupted with incorrect information. Our results unexpectedly show that Whisper may not understand the textual prompts in a human-expected way. Additionally, we find that performance improvement is not guaranteed even with stronger adherence to the topic information in textual prompts. It is also noted that English prompts generally outperform Mandarin ones on datasets of both languages, likely due to differences in training data distributions for these languages despite the mismatch with pre-training scenarios. Conversely, we discover that Whisper exhibits awareness of misleading information in language tokens by ignoring incorrect language tokens and focusing on the correct ones. In sum, We raise insightful questions about Whisper's prompt understanding and reveal its counter-intuitive behaviors. We encourage further studies.
title Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.05806