Dude, where's my utterance? Evaluating the effects of automatic segmentation and transcription on CPS detection
Fuente:
arXiv
Saved in:
| Main Authors: | Venkatesha, Videep, Bradford, Mariah, Blanchard, Nathaniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Exploration of Internal States in Collaborative Problem Solving
by: Anindho, Sifatul, et al.
Published: (2025)
by: Anindho, Sifatul, et al.
Published: (2025)
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
by: Palaskar, Shruti, et al.
Published: (2024)
by: Palaskar, Shruti, et al.
Published: (2024)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
by: Hui, Macarious, et al.
Published: (2024)
by: Hui, Macarious, et al.
Published: (2024)
Chord Colourizer: A Near Real-Time System for Visualizing Musical Key
by: Haimes, Paul
Published: (2025)
by: Haimes, Paul
Published: (2025)
Accessibility and Social Inclusivity: A Literature Review of Music Technology for Blind and Low Vision People
by: Zhang, Shumeng, et al.
Published: (2025)
by: Zhang, Shumeng, et al.
Published: (2025)
People are poorly equipped to detect AI-powered voice clones
by: Barrington, Sarah, et al.
Published: (2024)
by: Barrington, Sarah, et al.
Published: (2024)
Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice
by: Mandai, Yuto, et al.
Published: (2025)
by: Mandai, Yuto, et al.
Published: (2025)
Qualitative Approaches to Voice UX
by: Seaborn, Katie, et al.
Published: (2024)
by: Seaborn, Katie, et al.
Published: (2024)
Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
by: Fujii, Takao, et al.
Published: (2025)
by: Fujii, Takao, et al.
Published: (2025)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
Turn-taking annotation for quantitative and qualitative analyses of conversation
by: Kelterer, Anneliese, et al.
Published: (2025)
by: Kelterer, Anneliese, et al.
Published: (2025)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
by: Jeon, Hyunbae, et al.
Published: (2024)
by: Jeon, Hyunbae, et al.
Published: (2024)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
Are Expressions for Music Emotions the Same Across Cultures?
by: Celen, Elif, et al.
Published: (2025)
by: Celen, Elif, et al.
Published: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024)
by: Sharma, Roshan, et al.
Published: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
by: Kimelman, Robert G.
Published: (2024)
by: Kimelman, Robert G.
Published: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
by: Uro, Rémi, et al.
Published: (2024)
by: Uro, Rémi, et al.
Published: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
by: Fu, Yu-Kuan, et al.
Published: (2024)
by: Fu, Yu-Kuan, et al.
Published: (2024)
Examining Audio Communication Mechanisms for Supervising Fleets of Agricultural Robots
by: Kamboj, Abhi, et al.
Published: (2022)
by: Kamboj, Abhi, et al.
Published: (2022)
A Methodological Framework for Capturing Cognitive-Affective States in Collaborative Learning
by: Anindho, Sifatul, et al.
Published: (2025)
by: Anindho, Sifatul, et al.
Published: (2025)
Collecting Prosody in the Wild: A Content-Controlled, Privacy-First Smartphone Protocol and Empirical Evaluation
by: Koch, Timo K., et al.
Published: (2026)
by: Koch, Timo K., et al.
Published: (2026)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
by: Zheng, Naijun, et al.
Published: (2025)
by: Zheng, Naijun, et al.
Published: (2025)
Low-latency auditory spatial attention detection based on spectro-spatial features from EEG
by: Cai, Siqi, et al.
Published: (2021)
by: Cai, Siqi, et al.
Published: (2021)
Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
by: Rojas-Galeano, Sergio
Published: (2025)
by: Rojas-Galeano, Sergio
Published: (2025)
Beamforming-LLM: What, Where and When Did I Miss?
by: Choudhari, Vishal
Published: (2025)
by: Choudhari, Vishal
Published: (2025)
Towards Reliable Large Audio Language Model
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment
by: Yasuda, Yusuke, et al.
Published: (2024)
by: Yasuda, Yusuke, et al.
Published: (2024)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
by: Andrusenko, Andrei, et al.
Published: (2026)
by: Andrusenko, Andrei, et al.
Published: (2026)
More-than-Human Storytelling: Designing Longitudinal Narrative Engagements with Generative AI
by: Fabre, Émilie, et al.
Published: (2025)
by: Fabre, Émilie, et al.
Published: (2025)
Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
by: Dong, Xuefu, et al.
Published: (2025)
by: Dong, Xuefu, et al.
Published: (2025)
Cluster-to-Predict Affect Contours from Speech
by: Kuşçu, Gökhan, et al.
Published: (2024)
by: Kuşçu, Gökhan, et al.
Published: (2024)
Silent Speech Sentence Recognition with Six-Axis Accelerometers using Conformer and CTC Algorithm
by: Xie, Yudong, et al.
Published: (2025)
by: Xie, Yudong, et al.
Published: (2025)
Toward using Speech to Sense Student Emotion in Remote Learning Environments
by: Vyas, Sargam, et al.
Published: (2026)
by: Vyas, Sargam, et al.
Published: (2026)
Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs
by: Arya, Lalaram, et al.
Published: (2026)
by: Arya, Lalaram, et al.
Published: (2026)
The effect of self-motion and room familiarity on sound source localization in virtual environments
by: Isserstedt, Niklas, et al.
Published: (2024)
by: Isserstedt, Niklas, et al.
Published: (2024)
Evaluating Spatialized Auditory Cues for Rapid Attention Capture in XR
by: Kim, Yoonsang, et al.
Published: (2026)
by: Kim, Yoonsang, et al.
Published: (2026)
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
by: Kuhn, Korbinian, et al.
Published: (2025)
by: Kuhn, Korbinian, et al.
Published: (2025)
Similar Items
-
An Exploration of Internal States in Collaborative Problem Solving
by: Anindho, Sifatul, et al.
Published: (2025) -
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
by: Palaskar, Shruti, et al.
Published: (2024) -
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
by: Hui, Macarious, et al.
Published: (2024) -
Chord Colourizer: A Near Real-Time System for Visualizing Musical Key
by: Haimes, Paul
Published: (2025) -
Accessibility and Social Inclusivity: A Literature Review of Music Technology for Blind and Low Vision People
by: Zhang, Shumeng, et al.
Published: (2025)