Saved in:
| Main Authors: | Leite, Pedro H. L., Valadares, Pedro Benevenuto, Biscainho, Luiz W. P. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.30457 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution
by: Abreu, Wallace, et al.
Published: (2026)
by: Abreu, Wallace, et al.
Published: (2026)
Encoding of lexical tone in self-supervised models of spoken language
by: Shen, Gaofei, et al.
Published: (2024)
by: Shen, Gaofei, et al.
Published: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
by: Ogun, Sewade, et al.
Published: (2024)
by: Ogun, Sewade, et al.
Published: (2024)
Out-of-distribution generalisation in spoken language understanding
by: Porjazovski, Dejan, et al.
Published: (2024)
by: Porjazovski, Dejan, et al.
Published: (2024)
Transfer the linguistic representations from TTS to accent conversion with non-parallel data
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and state space models
by: Abreu, Wallace, et al.
Published: (2024)
by: Abreu, Wallace, et al.
Published: (2024)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
Is one brick enough to break the wall of spoken dialogue state tracking?
by: Druart, Lucas, et al.
Published: (2023)
by: Druart, Lucas, et al.
Published: (2023)
BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
by: Lavechin, Marvin, et al.
Published: (2023)
by: Lavechin, Marvin, et al.
Published: (2023)
CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
by: Carvalho, Carlos, et al.
Published: (2025)
by: Carvalho, Carlos, et al.
Published: (2025)
Optimizing the role of human evaluation in LLM-based spoken document summarization systems
by: Kroll, Margaret, et al.
Published: (2024)
by: Kroll, Margaret, et al.
Published: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
by: Uro, Rémi, et al.
Published: (2024)
by: Uro, Rémi, et al.
Published: (2024)
On the reliability of feature attribution methods for speech classification
by: Shen, Gaofei, et al.
Published: (2025)
by: Shen, Gaofei, et al.
Published: (2025)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
by: Corrêa, Pedro, et al.
Published: (2025)
by: Corrêa, Pedro, et al.
Published: (2025)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
by: Kamper, Herman, et al.
Published: (2025)
by: Kamper, Herman, et al.
Published: (2025)
MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR
by: Damianos, Dimitrios, et al.
Published: (2025)
by: Damianos, Dimitrios, et al.
Published: (2025)
MiLorE-SSL: Scaling Multilingual Capabilities in Self-Supervised Models without Forgetting
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Revisiting speech segmentation and lexicon learning with better features
by: Kamper, Herman, et al.
Published: (2024)
by: Kamper, Herman, et al.
Published: (2024)
Autoregressive Speech Synthesis without Vector Quantization
by: Meng, Lingwei, et al.
Published: (2024)
by: Meng, Lingwei, et al.
Published: (2024)
Lamer-SSL: Layer-aware Mixture of LoRA Experts for Continual Multilingual Expansion of Self-supervised Models without Forgetting
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Cross-lingual Alzheimer's Disease detection based on paralinguistic and pre-trained features
by: Chen, Xuchu, et al.
Published: (2023)
by: Chen, Xuchu, et al.
Published: (2023)
Speech foundation models in healthcare: Effect of layer selection on pathological speech feature prediction
by: Wiepert, Daniela A., et al.
Published: (2024)
by: Wiepert, Daniela A., et al.
Published: (2024)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
by: Prakash, Jeena, et al.
Published: (2025)
by: Prakash, Jeena, et al.
Published: (2025)
FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
Accent-VITS:accent transfer for end-to-end TTS
by: Ma, Linhan, et al.
Published: (2023)
by: Ma, Linhan, et al.
Published: (2023)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
by: Sivasankaran, Sunit, et al.
Published: (2024)
by: Sivasankaran, Sunit, et al.
Published: (2024)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
by: Xie, Jingran, et al.
Published: (2025)
by: Xie, Jingran, et al.
Published: (2025)
Bayesian Restoration of Audio Degraded by Low-Frequency Pulses Modeled via Gaussian Process
by: de Carvalho, Hugo Tremonte, et al.
Published: (2020)
by: de Carvalho, Hugo Tremonte, et al.
Published: (2020)
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
by: Prabhavalkar, Rohit, et al.
Published: (2024)
by: Prabhavalkar, Rohit, et al.
Published: (2024)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
by: Kim, Seung-Bin, et al.
Published: (2024)
by: Kim, Seung-Bin, et al.
Published: (2024)
Document Author Classification Using Parsed Language Structure
by: Moon, Todd K, et al.
Published: (2024)
by: Moon, Todd K, et al.
Published: (2024)
Preserving spoken content in voice anonymisation with character-level vocoder conditioning
by: Panariello, Michele, et al.
Published: (2024)
by: Panariello, Michele, et al.
Published: (2024)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
by: Singh, Jaskaran, et al.
Published: (2025)
by: Singh, Jaskaran, et al.
Published: (2025)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
by: Khaertdinov, Bulat, et al.
Published: (2024)
by: Khaertdinov, Bulat, et al.
Published: (2024)
Sylber 2.0: A Universal Syllable Embedding
by: Cho, Cheol Jun, et al.
Published: (2026)
by: Cho, Cheol Jun, et al.
Published: (2026)
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
by: Cho, Cheol Jun, et al.
Published: (2023)
by: Cho, Cheol Jun, et al.
Published: (2023)
Scaling Spoken Language Models with Syllabic Speech Tokenization
by: Lee, Nicholas, et al.
Published: (2025)
by: Lee, Nicholas, et al.
Published: (2025)
A multimodal LLM for the non-invasive decoding of spoken text from brain recordings
by: Hmamouche, Youssef, et al.
Published: (2024)
by: Hmamouche, Youssef, et al.
Published: (2024)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
by: Cho, Cheol Jun, et al.
Published: (2023)
by: Cho, Cheol Jun, et al.
Published: (2023)
Similar Items
-
FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution
by: Abreu, Wallace, et al.
Published: (2026) -
Encoding of lexical tone in self-supervised models of spoken language
by: Shen, Gaofei, et al.
Published: (2024) -
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
by: Ogun, Sewade, et al.
Published: (2024) -
Out-of-distribution generalisation in spoken language understanding
by: Porjazovski, Dejan, et al.
Published: (2024) -
Transfer the linguistic representations from TTS to accent conversion with non-parallel data
by: Chen, Xi, et al.
Published: (2024)