Classification of Spontaneous and Scripted Speech for Multilingual Audio
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Elisha, Shahar, McDowell, Andrew, Beguerisse-Díaz, Mariano, Benetos, Emmanouil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
SpeechTaxi: On Multilingual Semantic Speech Classification
von: Keller, Lennart, et al.
Veröffentlicht: (2024)
von: Keller, Lennart, et al.
Veröffentlicht: (2024)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
von: Yeo, Eunjung
Veröffentlicht: (2024)
von: Yeo, Eunjung
Veröffentlicht: (2024)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
von: Li, Song, et al.
Veröffentlicht: (2024)
von: Li, Song, et al.
Veröffentlicht: (2024)
Learning Music Audio Representations With Limited Data
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
Measuring Entrainment in Spontaneous Code-switched Speech
von: Bhattacharya, Debasmita, et al.
Veröffentlicht: (2023)
von: Bhattacharya, Debasmita, et al.
Veröffentlicht: (2023)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
CASPER: A Large Scale Spontaneous Speech Dataset
von: Xiao, Cihan, et al.
Veröffentlicht: (2025)
von: Xiao, Cihan, et al.
Veröffentlicht: (2025)
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
CAMEO: Collection of Multilingual Emotional Speech Corpora
von: Christop, Iwona, et al.
Veröffentlicht: (2025)
von: Christop, Iwona, et al.
Veröffentlicht: (2025)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
YODAS: Youtube-Oriented Dataset for Audio and Speech
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
von: Manakul, Potsawee, et al.
Veröffentlicht: (2025)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2025)
Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
von: Tao, Fuxiang, et al.
Veröffentlicht: (2024)
von: Tao, Fuxiang, et al.
Veröffentlicht: (2024)
LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech
von: Kim, Haechan, et al.
Veröffentlicht: (2024)
von: Kim, Haechan, et al.
Veröffentlicht: (2024)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Spontaneous Speech-Based Suicide Risk Detection
von: Gao, Yifan, et al.
Veröffentlicht: (2025)
von: Gao, Yifan, et al.
Veröffentlicht: (2025)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
Multilingual Stutter Event Detection for English, German, and Mandarin Speech
von: Haas, Felix, et al.
Veröffentlicht: (2026)
von: Haas, Felix, et al.
Veröffentlicht: (2026)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
What Are They Doing? Joint Audio-Speech Co-Reasoning
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
USAD: Universal Speech and Audio Representation via Distillation
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2025)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2025)
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
von: Piñeiro-Martín, Andrés, et al.
Veröffentlicht: (2024)
von: Piñeiro-Martín, Andrés, et al.
Veröffentlicht: (2024)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024) -
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025) -
SpeechTaxi: On Multilingual Semantic Speech Classification
von: Keller, Lennart, et al.
Veröffentlicht: (2024) -
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
von: Zhuo, Le, et al.
Veröffentlicht: (2023) -
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)