_version_ 1866915320609898496
author Tembine, Hamidou
Bamia, Issa
NDong, Massa
Coulibaly, Bakary
Traore, Oumar Issiaka
Traore, Moussa
Sanogo, Moussa
Sangare, Mamadou Eric
Kante, Salif
Yongueng, Daryl Noupa
Ali, Hafiz Tiomoko
Tiomoko, Malik
Laleye, Frejus
Djehiche, Boualem
Dipama, Wesmanegda Elisee
Saje, Idris Baba
Ibrahim, Hammid Mohammed
Sanogo, Moumini
Nininahazwe, Marie Coursel
Siita, Abdul-Latif
Mhlongo, Haine
Kouka, Teddy Nelvy Dieu Merci
Jeridi, Mariam Serine
Mupenge, Mutiyamuogo Parfait
Dehah, Lekoueiry
Bouko, Abdoul Aziz Bio Sidi
Zokoue, Wilfried Franceslas
Sambila, Odette Richette
Mbango, Alina RS
Diagouraga, Mady
Sanoussi, Oumarou Moussa
Dessalegn, Gizachew
Samoura, Mohamed Lamine
Coulibaly, Bintou Laetitia Audrey
author_facet Tembine, Hamidou
Bamia, Issa
NDong, Massa
Coulibaly, Bakary
Traore, Oumar Issiaka
Traore, Moussa
Sanogo, Moussa
Sangare, Mamadou Eric
Kante, Salif
Yongueng, Daryl Noupa
Ali, Hafiz Tiomoko
Tiomoko, Malik
Laleye, Frejus
Djehiche, Boualem
Dipama, Wesmanegda Elisee
Saje, Idris Baba
Ibrahim, Hammid Mohammed
Sanogo, Moumini
Nininahazwe, Marie Coursel
Siita, Abdul-Latif
Mhlongo, Haine
Kouka, Teddy Nelvy Dieu Merci
Jeridi, Mariam Serine
Mupenge, Mutiyamuogo Parfait
Dehah, Lekoueiry
Bouko, Abdoul Aziz Bio Sidi
Zokoue, Wilfried Franceslas
Sambila, Odette Richette
Mbango, Alina RS
Diagouraga, Mady
Sanoussi, Oumarou Moussa
Dessalegn, Gizachew
Samoura, Mohamed Lamine
Coulibaly, Bintou Laetitia Audrey
contents While global linguistic diversity spans more than 7164 recognized languages, the current dominant architecture of machine intelligence remains fundamentally biased toward written text. This bias excludes over 700 million people particularly in rural and remote regions who are audio-literate. In this work, we introduce a fully textless, audio-to-audio machine intelligence framework designed to serve this underserved population, and all the people who prefer audio-efficiency. Our contributions include novel Audio-to-Audio translation architectures that bypass text entirely, including spectrogram-, scalogram-, wavelet-, and unit-based models. Central to our approach is the Multiscale Audio-Semantic Transform (MAST), a representation that encodes tonal, prosodic, speaker, and expressive features. We further integrate MAST into a fractional diffusion of mean-field-type framework powered by fractional Brownian motion. It enables the generation of high-fidelity, semantically consistent speech without reliance on textual supervision. The result is a robust and scalable system capable of learning directly from raw audio, even in languages that are unwritten or rarely digitized. This work represents a fundamental shift toward audio-native machine intelligence systems, expanding access to language technologies for communities historically left out of the current machine intelligence ecosystem.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02443
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking the Barriers of Text-Hungry and Audio-Deficient AI
Tembine, Hamidou
Bamia, Issa
NDong, Massa
Coulibaly, Bakary
Traore, Oumar Issiaka
Traore, Moussa
Sanogo, Moussa
Sangare, Mamadou Eric
Kante, Salif
Yongueng, Daryl Noupa
Ali, Hafiz Tiomoko
Tiomoko, Malik
Laleye, Frejus
Djehiche, Boualem
Dipama, Wesmanegda Elisee
Saje, Idris Baba
Ibrahim, Hammid Mohammed
Sanogo, Moumini
Nininahazwe, Marie Coursel
Siita, Abdul-Latif
Mhlongo, Haine
Kouka, Teddy Nelvy Dieu Merci
Jeridi, Mariam Serine
Mupenge, Mutiyamuogo Parfait
Dehah, Lekoueiry
Bouko, Abdoul Aziz Bio Sidi
Zokoue, Wilfried Franceslas
Sambila, Odette Richette
Mbango, Alina RS
Diagouraga, Mady
Sanoussi, Oumarou Moussa
Dessalegn, Gizachew
Samoura, Mohamed Lamine
Coulibaly, Bintou Laetitia Audrey
Sound
Audio and Speech Processing
While global linguistic diversity spans more than 7164 recognized languages, the current dominant architecture of machine intelligence remains fundamentally biased toward written text. This bias excludes over 700 million people particularly in rural and remote regions who are audio-literate. In this work, we introduce a fully textless, audio-to-audio machine intelligence framework designed to serve this underserved population, and all the people who prefer audio-efficiency. Our contributions include novel Audio-to-Audio translation architectures that bypass text entirely, including spectrogram-, scalogram-, wavelet-, and unit-based models. Central to our approach is the Multiscale Audio-Semantic Transform (MAST), a representation that encodes tonal, prosodic, speaker, and expressive features. We further integrate MAST into a fractional diffusion of mean-field-type framework powered by fractional Brownian motion. It enables the generation of high-fidelity, semantically consistent speech without reliance on textual supervision. The result is a robust and scalable system capable of learning directly from raw audio, even in languages that are unwritten or rarely digitized. This work represents a fundamental shift toward audio-native machine intelligence systems, expanding access to language technologies for communities historically left out of the current machine intelligence ecosystem.
title Breaking the Barriers of Text-Hungry and Audio-Deficient AI
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.02443