Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911963063255040 |
|---|---|
| author | Saeed, Muhammad Saad Nawaz, Shah Tahir, Muhammad Salman Das, Rohan Kumar Zaheer, Muhammad Zaigham Moscati, Marta Schedl, Markus Khan, Muhammad Haris Nandakumar, Karthik Yousaf, Muhammad Haroon |
| author_facet | Saeed, Muhammad Saad Nawaz, Shah Tahir, Muhammad Salman Das, Rohan Kumar Zaheer, Muhammad Zaigham Moscati, Marta Schedl, Markus Khan, Muhammad Haris Nandakumar, Karthik Yousaf, Muhammad Haroon |
| contents | The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associating face and voice of a person has gained attention due to presence of unique correlation between them. The Face-voice Association in Multilingual Environments (FAME) Challenge 2024 focuses on exploring face-voice association under a unique condition of multilingual scenario. This condition is inspired from the fact that half of the world's population is bilingual and most often people communicate under multilingual scenario. The challenge uses a dataset namely, Multilingual Audio-Visual (MAV-Celeb) for exploring face-voice association in multilingual environments. This report provides the details of the challenge, dataset, baselines and task details for the FAME Challenge. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_09342 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan Saeed, Muhammad Saad Nawaz, Shah Tahir, Muhammad Salman Das, Rohan Kumar Zaheer, Muhammad Zaigham Moscati, Marta Schedl, Markus Khan, Muhammad Haris Nandakumar, Karthik Yousaf, Muhammad Haroon Computer Vision and Pattern Recognition Sound Audio and Speech Processing The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associating face and voice of a person has gained attention due to presence of unique correlation between them. The Face-voice Association in Multilingual Environments (FAME) Challenge 2024 focuses on exploring face-voice association under a unique condition of multilingual scenario. This condition is inspired from the fact that half of the world's population is bilingual and most often people communicate under multilingual scenario. The challenge uses a dataset namely, Multilingual Audio-Visual (MAV-Celeb) for exploring face-voice association in multilingual environments. This report provides the details of the challenge, dataset, baselines and task details for the FAME Challenge. |
| title | Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan |
| topic | Computer Vision and Pattern Recognition Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2404.09342 |