Saved in:
| Main Authors: | Jin, Liuyi, Gunawardena, Pasan, Haroon, Amran, Wang, Runzhi, Lee, Sangwoo, Stoleru, Radu, Middleton, Michael, Huo, Zepeng, Kim, Jeeeun, Moats, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.13078 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services
by: Jin, Liuyi, et al.
Published: (2025)
by: Jin, Liuyi, et al.
Published: (2025)
Blind Identification of Binaural Room Impulse Responses from Smart Glasses
by: Deppisch, Thomas, et al.
Published: (2024)
by: Deppisch, Thomas, et al.
Published: (2024)
MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Multitask Learning with Capsule Networks for Speech-to-Intent Applications
by: Poncelet, Jakob, et al.
Published: (2020)
by: Poncelet, Jakob, et al.
Published: (2020)
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
by: Yang, Yufeng, et al.
Published: (2025)
by: Yang, Yufeng, et al.
Published: (2025)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
by: Yang, Yufeng, et al.
Published: (2024)
by: Yang, Yufeng, et al.
Published: (2024)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
by: Xu, Zhongweiyang, et al.
Published: (2024)
by: Xu, Zhongweiyang, et al.
Published: (2024)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
by: Shao, Mingchen, et al.
Published: (2025)
by: Shao, Mingchen, et al.
Published: (2025)
Identifying and Calibrating Overconfidence in Noisy Speech Recognition
by: Huo, Mingyue, et al.
Published: (2025)
by: Huo, Mingyue, et al.
Published: (2025)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
by: Hao, Xiang, et al.
Published: (2020)
by: Hao, Xiang, et al.
Published: (2020)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
A New Time Series Similarity Measure and Its Smart Grid Applications
by: Yuan, Rui, et al.
Published: (2023)
by: Yuan, Rui, et al.
Published: (2023)
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
by: Kumar, Shashi, et al.
Published: (2025)
by: Kumar, Shashi, et al.
Published: (2025)
ReHear: Iterative Pseudo-Label Refinement for Semi-Supervised Speech Recognition via Audio Large Language Models
by: Liu, Zefang, et al.
Published: (2026)
by: Liu, Zefang, et al.
Published: (2026)
A Multimodal Framework for the Assessment of the Schizophrenia Spectrum
by: Premananth, Gowtham, et al.
Published: (2024)
by: Premananth, Gowtham, et al.
Published: (2024)
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
by: Wu, Ke, et al.
Published: (2026)
by: Wu, Ke, et al.
Published: (2026)
VoiceVector: Multimodal Enrolment Vectors for Speaker Separation
by: Rahimi, Akam, et al.
Published: (2025)
by: Rahimi, Akam, et al.
Published: (2025)
A Survey of Audio Reasoning in Multimodal Foundation Models
by: Guo, Zhihan, et al.
Published: (2026)
by: Guo, Zhihan, et al.
Published: (2026)
Are Multimodal Foundation Models All That Is Needed for Emofake Detection?
by: Akhtar, Mohd Mujtaba, et al.
Published: (2025)
by: Akhtar, Mohd Mujtaba, et al.
Published: (2025)
Towards Multimodal Query-Based Spatial Audio Source Extraction
by: Yu, Chenxin, et al.
Published: (2025)
by: Yu, Chenxin, et al.
Published: (2025)
VoxATtack: A Multimodal Attack on Voice Anonymization Systems
by: Aloradi, Ahmad, et al.
Published: (2025)
by: Aloradi, Ahmad, et al.
Published: (2025)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
by: Baroudi, Séverin, et al.
Published: (2026)
by: Baroudi, Séverin, et al.
Published: (2026)
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
by: Chen, Yafeng, et al.
Published: (2024)
by: Chen, Yafeng, et al.
Published: (2024)
Multitask Learning for Grapheme-to-Phoneme Conversion of Anglicisms in German Speech Recognition
by: Pritzen, Julia, et al.
Published: (2021)
by: Pritzen, Julia, et al.
Published: (2021)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
by: Zhang, Wenyu, et al.
Published: (2025)
by: Zhang, Wenyu, et al.
Published: (2025)
Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
by: Hsieh, Tsun-An, et al.
Published: (2024)
by: Hsieh, Tsun-An, et al.
Published: (2024)
Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
by: Wang, Siyin, et al.
Published: (2025)
by: Wang, Siyin, et al.
Published: (2025)
Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs
by: Langman, Ryan, et al.
Published: (2024)
by: Langman, Ryan, et al.
Published: (2024)
Leveraging Cascaded Binary Classification and Multimodal Fusion for Dementia Detection through Spontaneous Speech
by: Liu, Yin-Long, et al.
Published: (2025)
by: Liu, Yin-Long, et al.
Published: (2025)
Predicting Cognitive Decline: A Multimodal AI Approach to Dementia Screening from Speech
by: Chi, Lei, et al.
Published: (2025)
by: Chi, Lei, et al.
Published: (2025)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
by: Zhang, Jing-Xuan, et al.
Published: (2025)
by: Zhang, Jing-Xuan, et al.
Published: (2025)
Multimodal Deep Learning Method for Real-Time Spatial Room Impulse Response Computing
by: Li, Zhiyu, et al.
Published: (2026)
by: Li, Zhiyu, et al.
Published: (2026)
Empowering Multimodal Respiratory Sound Classification with Counterfactual Adversarial Debiasing for Out-of-Distribution Robustness
by: Koo, Heejoon, et al.
Published: (2025)
by: Koo, Heejoon, et al.
Published: (2025)
Adaptable Non-parametric Approach for Speech-based Symptom Assessment: Isolating Private Medical Data in a Retrieval Datastore
by: Chen, Yu-Wen, et al.
Published: (2025)
by: Chen, Yu-Wen, et al.
Published: (2025)
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
by: Huo, Yanru, et al.
Published: (2025)
by: Huo, Yanru, et al.
Published: (2025)
HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart Speakers
by: Xie, Yadong, et al.
Published: (2025)
by: Xie, Yadong, et al.
Published: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
by: Yang, Runyan, et al.
Published: (2024)
by: Yang, Runyan, et al.
Published: (2024)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
by: Huo, Mingyue, et al.
Published: (2025)
by: Huo, Mingyue, et al.
Published: (2025)
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
by: Langman, Ryan, et al.
Published: (2025)
by: Langman, Ryan, et al.
Published: (2025)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
by: Eliav, Amit, et al.
Published: (2024)
by: Eliav, Amit, et al.
Published: (2024)
Similar Items
-
Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services
by: Jin, Liuyi, et al.
Published: (2025) -
Blind Identification of Binaural Room Impulse Responses from Smart Glasses
by: Deppisch, Thomas, et al.
Published: (2024) -
MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses
by: Liu, Yang, et al.
Published: (2025) -
Multitask Learning with Capsule Networks for Speech-to-Intent Applications
by: Poncelet, Jakob, et al.
Published: (2020) -
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
by: Yang, Yufeng, et al.
Published: (2025)