Navigating Speech Recording Collections with AI-Generated Illustrations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Håland, Sirina, Strøm, Trond Karlsen, Galuščáková, Petra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The language of sound search: Examining User Queries in Audio Search Engines
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2024)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2024)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
Are Expressions for Music Emotions the Same Across Cultures?
von: Celen, Elif, et al.
Veröffentlicht: (2025)
von: Celen, Elif, et al.
Veröffentlicht: (2025)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
von: Kimelman, Robert G.
Veröffentlicht: (2024)
von: Kimelman, Robert G.
Veröffentlicht: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
A Framework for AI assisted Musical Devices
von: Civit, Miguel, et al.
Veröffentlicht: (2024)
von: Civit, Miguel, et al.
Veröffentlicht: (2024)
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
von: Gomez, Frank Palma, et al.
Veröffentlicht: (2024)
von: Gomez, Frank Palma, et al.
Veröffentlicht: (2024)
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
A GEN AI Framework for Medical Note Generation
von: Leong, Hui Yi, et al.
Veröffentlicht: (2024)
von: Leong, Hui Yi, et al.
Veröffentlicht: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The language of sound search: Examining User Queries in Audio Search Engines
von: Weck, Benno, et al.
Veröffentlicht: (2024) -
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024) -
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024) -
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025) -
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2024)