Spontaneous Informal Speech Dataset for Punctuation Restoration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xing Yi, Beigi, Homayoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
von: Liu, Xing Yi, et al.
Veröffentlicht: (2023)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2023)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
LLAMAPIE: Proactive In-Ear Conversation Assistants
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
Literary and Colloquial Tamil Dialect Identification
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
A conversational gesture synthesis system based on emotions and semantics
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
von: Liu, Tianyun
Veröffentlicht: (2025)
von: Liu, Tianyun
Veröffentlicht: (2025)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
von: Echeverria, Alex, et al.
Veröffentlicht: (2025)
von: Echeverria, Alex, et al.
Veröffentlicht: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
von: Jaiswal, Pranay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Pranay, et al.
Veröffentlicht: (2024)
DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition
von: Kim, Hyunju, et al.
Veröffentlicht: (2024)
von: Kim, Hyunju, et al.
Veröffentlicht: (2024)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
von: Doukhan, David, et al.
Veröffentlicht: (2024)
von: Doukhan, David, et al.
Veröffentlicht: (2024)
Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
von: Garcia, Nelly, et al.
Veröffentlicht: (2026)
von: Garcia, Nelly, et al.
Veröffentlicht: (2026)
Improving AI-generated music with user-guided training
von: Singh, Vishwa Mohan, et al.
Veröffentlicht: (2025)
von: Singh, Vishwa Mohan, et al.
Veröffentlicht: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Multimodal Contextualized Semantic Parsing from Speech
von: Voas, Jordan, et al.
Veröffentlicht: (2024)
von: Voas, Jordan, et al.
Veröffentlicht: (2024)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Towards Decoding Brain Activity During Passive Listening of Speech
von: Fodor, Milán András, et al.
Veröffentlicht: (2024)
von: Fodor, Milán András, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
von: Liu, Xing Yi, et al.
Veröffentlicht: (2023) -
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026) -
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025) -
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025) -
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)