Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Saif, A F M, Chen, Lisha, Cui, Xiaodong, Lu, Songtao, Kingsbury, Brian, Chen, Tianyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
by: Jiang, Liuyuan, et al.
Published: (2025)
by: Jiang, Liuyuan, et al.
Published: (2025)
Stable Acoustic Relay Assignment with High Throughput via Lase Chaos-based Reinforcement Learning
by: Chen, Zengjing, et al.
Published: (2025)
by: Chen, Zengjing, et al.
Published: (2025)
Multi-Source Localization and Data Association for Time-Difference of Arrival Measurements
by: Flood, Gabrielle, et al.
Published: (2024)
by: Flood, Gabrielle, et al.
Published: (2024)
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
by: Lovelace, Justin, et al.
Published: (2025)
by: Lovelace, Justin, et al.
Published: (2025)
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
by: Gao, Xiaoxue, et al.
Published: (2026)
by: Gao, Xiaoxue, et al.
Published: (2026)
Optimized Loudspeaker Panning for Adaptive Sound-Field Correction and Non-stationary Listening Areas
by: Luo, Yuancheng
Published: (2025)
by: Luo, Yuancheng
Published: (2025)
Constraint Optimized Multichannel Mixer-limiter Design
by: Luo, Yuancheng, et al.
Published: (2025)
by: Luo, Yuancheng, et al.
Published: (2025)
Time Segmented Beamforming via Dynamic Programming: Theory and Implementation
by: Mittal, Manan, et al.
Published: (2026)
by: Mittal, Manan, et al.
Published: (2026)
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
by: Zezario, Ryandhimas E., et al.
Published: (2021)
by: Zezario, Ryandhimas E., et al.
Published: (2021)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
by: Zezario, Ryandhimas E., et al.
Published: (2023)
by: Zezario, Ryandhimas E., et al.
Published: (2023)
Morphogenesis of sound creates acoustic rainbows
by: Christiansen, Rasmus E., et al.
Published: (2024)
by: Christiansen, Rasmus E., et al.
Published: (2024)
Benchmarking Humans and Machines on Complex Multilingual Speech Understanding Tasks
by: Kankanala, Sai Samrat, et al.
Published: (2025)
by: Kankanala, Sai Samrat, et al.
Published: (2025)
FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning
by: Chen, Lisha, et al.
Published: (2024)
by: Chen, Lisha, et al.
Published: (2024)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
by: Fan, Xulin, et al.
Published: (2026)
by: Fan, Xulin, et al.
Published: (2026)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
by: Zhou, Yixuan, et al.
Published: (2024)
by: Zhou, Yixuan, et al.
Published: (2024)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
by: Fu, Yonggan, et al.
Published: (2022)
by: Fu, Yonggan, et al.
Published: (2022)
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
by: Hirschkind, Nameer, et al.
Published: (2024)
by: Hirschkind, Nameer, et al.
Published: (2024)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
by: Chen, William, et al.
Published: (2025)
by: Chen, William, et al.
Published: (2025)
Integrating Posture Control in Speech Motor Models: A Parallel-Structured Simulation Approach
by: Liu, Yadong, et al.
Published: (2024)
by: Liu, Yadong, et al.
Published: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
by: Ulgen, Ismail Rasim, et al.
Published: (2025)
by: Ulgen, Ismail Rasim, et al.
Published: (2025)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
by: Mu, Bingshen, et al.
Published: (2025)
by: Mu, Bingshen, et al.
Published: (2025)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
by: Narain, Jaya, et al.
Published: (2025)
by: Narain, Jaya, et al.
Published: (2025)
Towards Cross-Task Suicide Risk Detection via Speech LLM
by: Li, Jialun, et al.
Published: (2025)
by: Li, Jialun, et al.
Published: (2025)
Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
by: Khorrami, Khazar, et al.
Published: (2023)
by: Khorrami, Khazar, et al.
Published: (2023)
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
by: Beyene, Luel Hagos, et al.
Published: (2025)
by: Beyene, Luel Hagos, et al.
Published: (2025)
SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
by: Chang, Kai-Wei, et al.
Published: (2024)
by: Chang, Kai-Wei, et al.
Published: (2024)
Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
by: Serrand, Coralie, et al.
Published: (2025)
by: Serrand, Coralie, et al.
Published: (2025)
Masked Audio Modeling with CLAP and Multi-Objective Learning
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
by: Sarkar, Eklavya, et al.
Published: (2025)
by: Sarkar, Eklavya, et al.
Published: (2025)
Personal Sound Zones and Shielded Localized Communication through Active Acoustic Control
by: Egarguin, Neil Jerome A., et al.
Published: (2024)
by: Egarguin, Neil Jerome A., et al.
Published: (2024)
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
by: Park, Joonyong, et al.
Published: (2025)
by: Park, Joonyong, et al.
Published: (2025)
From Black Box to Biomarker: Sparse Autoencoders for Interpreting Speech Models of Parkinson's Disease
by: Plantinga, Peter, et al.
Published: (2025)
by: Plantinga, Peter, et al.
Published: (2025)
Noise-aware Speech Enhancement using Diffusion Probabilistic Model
by: Hu, Yuchen, et al.
Published: (2023)
by: Hu, Yuchen, et al.
Published: (2023)
Exploring the limits of decoder-only models trained on public speech recognition corpora
by: Gupta, Ankit, et al.
Published: (2024)
by: Gupta, Ankit, et al.
Published: (2024)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
by: Mahapatra, Aurosweta, et al.
Published: (2025)
by: Mahapatra, Aurosweta, et al.
Published: (2025)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
by: Ji, Shengpeng, et al.
Published: (2023)
by: Ji, Shengpeng, et al.
Published: (2023)
Open-Source System for Multilingual Translation and Cloned Speech Synthesis
by: Cámara, Mateo, et al.
Published: (2025)
by: Cámara, Mateo, et al.
Published: (2025)
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
by: Gao, Xiaoxue, et al.
Published: (2025)
by: Gao, Xiaoxue, et al.
Published: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
by: Liao, Shijia, et al.
Published: (2024)
by: Liao, Shijia, et al.
Published: (2024)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
by: Wang, Yuancheng, et al.
Published: (2025)
by: Wang, Yuancheng, et al.
Published: (2025)
Similar Items
-
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
by: Jiang, Liuyuan, et al.
Published: (2025) -
Stable Acoustic Relay Assignment with High Throughput via Lase Chaos-based Reinforcement Learning
by: Chen, Zengjing, et al.
Published: (2025) -
Multi-Source Localization and Data Association for Time-Difference of Arrival Measurements
by: Flood, Gabrielle, et al.
Published: (2024) -
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
by: Lovelace, Justin, et al.
Published: (2025) -
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
by: Gao, Xiaoxue, et al.
Published: (2026)