Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Bin, Zou, Xunlong, Sun, Shuo, Zhang, Wenyu, He, Yingxu, Liu, Zhuohan, Wei, Chengwei, Chen, Nancy F., Aw, AiTi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AudioBench: A Universal Benchmark for Audio Large Language Models
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
Bridging Language Gaps in Audio-Text Retrieval
von: Yan, Zhiyong, et al.
Veröffentlicht: (2024)
von: Yan, Zhiyong, et al.
Veröffentlicht: (2024)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
SYKI-SVC: Advancing Singing Voice Conversion with Post-Processing Innovations and an Open-Source Professional Testset
von: Zhou, Yiquan, et al.
Veröffentlicht: (2025)
von: Zhou, Yiquan, et al.
Veröffentlicht: (2025)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
von: Hu, Jing, et al.
Veröffentlicht: (2026)
von: Hu, Jing, et al.
Veröffentlicht: (2026)
WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
von: Karystinaios, Emmanouil
Veröffentlicht: (2025)
von: Karystinaios, Emmanouil
Veröffentlicht: (2025)
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
QuarkAudio Technical Report
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
DARAS: Dynamic Audio-Room Acoustic Synthesis for Blind Room Impulse Response Estimation
von: Wang, Chunxi, et al.
Veröffentlicht: (2025)
von: Wang, Chunxi, et al.
Veröffentlicht: (2025)
Mitigating Category Imbalance: Fosafer System for the Multimodal Emotion and Intent Joint Understanding Challenge
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
A Multi-stage Low-latency Enhancement System for Hearing Aids
von: Ouyang, Chengwei, et al.
Veröffentlicht: (2025)
von: Ouyang, Chengwei, et al.
Veröffentlicht: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
Bridge-SR: Schrödinger Bridge for Efficient SR
von: Li, Chang, et al.
Veröffentlicht: (2025)
von: Li, Chang, et al.
Veröffentlicht: (2025)
Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals
von: Chen, Zhang, et al.
Veröffentlicht: (2026)
von: Chen, Zhang, et al.
Veröffentlicht: (2026)
The Florence Price Art Song Dataset and Piano Accompaniment Generator
von: He, Tao-Tao, et al.
Veröffentlicht: (2025)
von: He, Tao-Tao, et al.
Veröffentlicht: (2025)
LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition
von: He, Zhining, et al.
Veröffentlicht: (2025)
von: He, Zhining, et al.
Veröffentlicht: (2025)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
A Hybrid Discriminative and Generative System for Universal Speech Enhancement
von: Liu, Yinghao, et al.
Veröffentlicht: (2026)
von: Liu, Yinghao, et al.
Veröffentlicht: (2026)
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
von: Yin, Jie, et al.
Veröffentlicht: (2024)
von: Yin, Jie, et al.
Veröffentlicht: (2024)
SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
von: Zhao, Changjiang, et al.
Veröffentlicht: (2024)
von: Zhao, Changjiang, et al.
Veröffentlicht: (2024)
SS-BRPE: Self-Supervised Blind Room Parameter Estimation Using Attention Mechanisms
von: Wang, Chunxi, et al.
Veröffentlicht: (2024)
von: Wang, Chunxi, et al.
Veröffentlicht: (2024)
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
Advances in Speech Separation: Techniques, Challenges, and Future Trends
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
Can Large Language Models Understand Spatial Audio?
von: Tang, Changli, et al.
Veröffentlicht: (2024)
von: Tang, Changli, et al.
Veröffentlicht: (2024)
FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
von: Jiang, Han, et al.
Veröffentlicht: (2024)
von: Jiang, Han, et al.
Veröffentlicht: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AudioBench: A Universal Benchmark for Audio Large Language Models
von: Wang, Bin, et al.
Veröffentlicht: (2024) -
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024) -
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025) -
ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025) -
Bridging Language Gaps in Audio-Text Retrieval
von: Yan, Zhiyong, et al.
Veröffentlicht: (2024)