Breaking the Barriers of Text-Hungry and Audio-Deficient AI
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915320609898496 |
|---|---|
| author | Tembine, Hamidou Bamia, Issa NDong, Massa Coulibaly, Bakary Traore, Oumar Issiaka Traore, Moussa Sanogo, Moussa Sangare, Mamadou Eric Kante, Salif Yongueng, Daryl Noupa Ali, Hafiz Tiomoko Tiomoko, Malik Laleye, Frejus Djehiche, Boualem Dipama, Wesmanegda Elisee Saje, Idris Baba Ibrahim, Hammid Mohammed Sanogo, Moumini Nininahazwe, Marie Coursel Siita, Abdul-Latif Mhlongo, Haine Kouka, Teddy Nelvy Dieu Merci Jeridi, Mariam Serine Mupenge, Mutiyamuogo Parfait Dehah, Lekoueiry Bouko, Abdoul Aziz Bio Sidi Zokoue, Wilfried Franceslas Sambila, Odette Richette Mbango, Alina RS Diagouraga, Mady Sanoussi, Oumarou Moussa Dessalegn, Gizachew Samoura, Mohamed Lamine Coulibaly, Bintou Laetitia Audrey |
| author_facet | Tembine, Hamidou Bamia, Issa NDong, Massa Coulibaly, Bakary Traore, Oumar Issiaka Traore, Moussa Sanogo, Moussa Sangare, Mamadou Eric Kante, Salif Yongueng, Daryl Noupa Ali, Hafiz Tiomoko Tiomoko, Malik Laleye, Frejus Djehiche, Boualem Dipama, Wesmanegda Elisee Saje, Idris Baba Ibrahim, Hammid Mohammed Sanogo, Moumini Nininahazwe, Marie Coursel Siita, Abdul-Latif Mhlongo, Haine Kouka, Teddy Nelvy Dieu Merci Jeridi, Mariam Serine Mupenge, Mutiyamuogo Parfait Dehah, Lekoueiry Bouko, Abdoul Aziz Bio Sidi Zokoue, Wilfried Franceslas Sambila, Odette Richette Mbango, Alina RS Diagouraga, Mady Sanoussi, Oumarou Moussa Dessalegn, Gizachew Samoura, Mohamed Lamine Coulibaly, Bintou Laetitia Audrey |
| contents | While global linguistic diversity spans more than 7164 recognized languages, the current dominant architecture of machine intelligence remains fundamentally biased toward written text. This bias excludes over 700 million people particularly in rural and remote regions who are audio-literate. In this work, we introduce a fully textless, audio-to-audio machine intelligence framework designed to serve this underserved population, and all the people who prefer audio-efficiency. Our contributions include novel Audio-to-Audio translation architectures that bypass text entirely, including spectrogram-, scalogram-, wavelet-, and unit-based models. Central to our approach is the Multiscale Audio-Semantic Transform (MAST), a representation that encodes tonal, prosodic, speaker, and expressive features. We further integrate MAST into a fractional diffusion of mean-field-type framework powered by fractional Brownian motion. It enables the generation of high-fidelity, semantically consistent speech without reliance on textual supervision. The result is a robust and scalable system capable of learning directly from raw audio, even in languages that are unwritten or rarely digitized. This work represents a fundamental shift toward audio-native machine intelligence systems, expanding access to language technologies for communities historically left out of the current machine intelligence ecosystem. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_02443 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Breaking the Barriers of Text-Hungry and Audio-Deficient AI Tembine, Hamidou Bamia, Issa NDong, Massa Coulibaly, Bakary Traore, Oumar Issiaka Traore, Moussa Sanogo, Moussa Sangare, Mamadou Eric Kante, Salif Yongueng, Daryl Noupa Ali, Hafiz Tiomoko Tiomoko, Malik Laleye, Frejus Djehiche, Boualem Dipama, Wesmanegda Elisee Saje, Idris Baba Ibrahim, Hammid Mohammed Sanogo, Moumini Nininahazwe, Marie Coursel Siita, Abdul-Latif Mhlongo, Haine Kouka, Teddy Nelvy Dieu Merci Jeridi, Mariam Serine Mupenge, Mutiyamuogo Parfait Dehah, Lekoueiry Bouko, Abdoul Aziz Bio Sidi Zokoue, Wilfried Franceslas Sambila, Odette Richette Mbango, Alina RS Diagouraga, Mady Sanoussi, Oumarou Moussa Dessalegn, Gizachew Samoura, Mohamed Lamine Coulibaly, Bintou Laetitia Audrey Sound Audio and Speech Processing While global linguistic diversity spans more than 7164 recognized languages, the current dominant architecture of machine intelligence remains fundamentally biased toward written text. This bias excludes over 700 million people particularly in rural and remote regions who are audio-literate. In this work, we introduce a fully textless, audio-to-audio machine intelligence framework designed to serve this underserved population, and all the people who prefer audio-efficiency. Our contributions include novel Audio-to-Audio translation architectures that bypass text entirely, including spectrogram-, scalogram-, wavelet-, and unit-based models. Central to our approach is the Multiscale Audio-Semantic Transform (MAST), a representation that encodes tonal, prosodic, speaker, and expressive features. We further integrate MAST into a fractional diffusion of mean-field-type framework powered by fractional Brownian motion. It enables the generation of high-fidelity, semantically consistent speech without reliance on textual supervision. The result is a robust and scalable system capable of learning directly from raw audio, even in languages that are unwritten or rarely digitized. This work represents a fundamental shift toward audio-native machine intelligence systems, expanding access to language technologies for communities historically left out of the current machine intelligence ecosystem. |
| title | Breaking the Barriers of Text-Hungry and Audio-Deficient AI |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.02443 |