Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929520381001728 |
|---|---|
| author | Chen, Chen Li, Xiaolou Liu, Zehua Li, Lantian Wang, Dong |
| author_facet | Chen, Chen Li, Xiaolou Liu, Zehua Li, Lantian Wang, Dong |
| contents | In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip reading, audio-visual speech recognition, and visual-to-speech synthesis. Although significant success has been achieved, theoretical analysis is still insufficient for audio-visual tasks. This paper presents a quantitative analysis based on information theory, focusing on information intersection between different modalities. Our results show that this analysis is valuable for understanding the difficulties of audio-visual processing tasks as well as the benefits that could be obtained by modality integration. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_19575 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective Chen, Chen Li, Xiaolou Liu, Zehua Li, Lantian Wang, Dong Sound Computation and Language Multimedia Audio and Speech Processing In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip reading, audio-visual speech recognition, and visual-to-speech synthesis. Although significant success has been achieved, theoretical analysis is still insufficient for audio-visual tasks. This paper presents a quantitative analysis based on information theory, focusing on information intersection between different modalities. Our results show that this analysis is valuable for understanding the difficulties of audio-visual processing tasks as well as the benefits that could be obtained by modality integration. |
| title | Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective |
| topic | Sound Computation and Language Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.19575 |