Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment
Fuente:
arXiv
Saved in:
| Main Authors: | Yasuda, Yusuke, Toda, Tomoki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A conversational gesture synthesis system based on emotions and semantics
by: Hoang-Minh, Thanh
Published: (2025)
by: Hoang-Minh, Thanh
Published: (2025)
Spontaneous Informal Speech Dataset for Punctuation Restoration
by: Liu, Xing Yi, et al.
Published: (2024)
by: Liu, Xing Yi, et al.
Published: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
by: Kawamura, Kazuki, et al.
Published: (2024)
by: Kawamura, Kazuki, et al.
Published: (2024)
Literary and Colloquial Tamil Dialect Identification
by: Nanmalar, M., et al.
Published: (2024)
by: Nanmalar, M., et al.
Published: (2024)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
by: Torgashov, Nikita, et al.
Published: (2025)
by: Torgashov, Nikita, et al.
Published: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience
by: Chang, Andrew, et al.
Published: (2025)
by: Chang, Andrew, et al.
Published: (2025)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
by: Zhang, Yixiao, et al.
Published: (2023)
by: Zhang, Yixiao, et al.
Published: (2023)
LLAMAPIE: Proactive In-Ear Conversation Assistants
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
by: Torgashov, Nikita, et al.
Published: (2026)
by: Torgashov, Nikita, et al.
Published: (2026)
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
by: Liu, Xing Yi, et al.
Published: (2023)
by: Liu, Xing Yi, et al.
Published: (2023)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
by: Doukhan, David, et al.
Published: (2024)
by: Doukhan, David, et al.
Published: (2024)
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
by: Kumar, Satyam, et al.
Published: (2024)
by: Kumar, Satyam, et al.
Published: (2024)
Open-Source Conversational AI with SpeechBrain 1.0
by: Ravanelli, Mirco, et al.
Published: (2024)
by: Ravanelli, Mirco, et al.
Published: (2024)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
by: Wang, Xihuai, et al.
Published: (2025)
by: Wang, Xihuai, et al.
Published: (2025)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
by: Liu, Tianyun
Published: (2025)
by: Liu, Tianyun
Published: (2025)
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
by: Palaskar, Shruti, et al.
Published: (2024)
by: Palaskar, Shruti, et al.
Published: (2024)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
by: Ali, Hasmot, et al.
Published: (2024)
by: Ali, Hasmot, et al.
Published: (2024)
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
by: Jaiswal, Pranay, et al.
Published: (2024)
by: Jaiswal, Pranay, et al.
Published: (2024)
DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition
by: Kim, Hyunju, et al.
Published: (2024)
by: Kim, Hyunju, et al.
Published: (2024)
Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
by: Garcia, Nelly, et al.
Published: (2026)
by: Garcia, Nelly, et al.
Published: (2026)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
by: Yuan, Kuang, et al.
Published: (2025)
by: Yuan, Kuang, et al.
Published: (2025)
Improving AI-generated music with user-guided training
by: Singh, Vishwa Mohan, et al.
Published: (2025)
by: Singh, Vishwa Mohan, et al.
Published: (2025)
Human Feedback Driven Dynamic Speech Emotion Recognition
by: Fedorov, Ilya, et al.
Published: (2025)
by: Fedorov, Ilya, et al.
Published: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
by: Sankey-Olsen, Cuno, et al.
Published: (2025)
by: Sankey-Olsen, Cuno, et al.
Published: (2025)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
by: Xie, Zhifei, et al.
Published: (2024)
by: Xie, Zhifei, et al.
Published: (2024)
The language of sound search: Examining User Queries in Audio Search Engines
by: Weck, Benno, et al.
Published: (2024)
by: Weck, Benno, et al.
Published: (2024)
Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
by: Echeverria, Alex, et al.
Published: (2025)
by: Echeverria, Alex, et al.
Published: (2025)
EmoHeal: An End-to-End System for Personalized Therapeutic Music Retrieval from Fine-grained Emotions
by: Wan, Xinchen, et al.
Published: (2025)
by: Wan, Xinchen, et al.
Published: (2025)
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication
by: Nakilcioglu, Emin Cagatay, et al.
Published: (2023)
by: Nakilcioglu, Emin Cagatay, et al.
Published: (2023)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
by: Jeon, Hyunbae, et al.
Published: (2024)
by: Jeon, Hyunbae, et al.
Published: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024)
by: Sharma, Roshan, et al.
Published: (2024)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
by: Hui, Macarious, et al.
Published: (2024)
by: Hui, Macarious, et al.
Published: (2024)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
by: Kimelman, Robert G.
Published: (2024)
by: Kimelman, Robert G.
Published: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
by: Uro, Rémi, et al.
Published: (2024)
by: Uro, Rémi, et al.
Published: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
by: Fu, Yu-Kuan, et al.
Published: (2024)
by: Fu, Yu-Kuan, et al.
Published: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
Similar Items
-
A conversational gesture synthesis system based on emotions and semantics
by: Hoang-Minh, Thanh
Published: (2025) -
Spontaneous Informal Speech Dataset for Punctuation Restoration
by: Liu, Xing Yi, et al.
Published: (2024) -
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
by: Kawamura, Kazuki, et al.
Published: (2024) -
Literary and Colloquial Tamil Dialect Identification
by: Nanmalar, M., et al.
Published: (2024) -
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
by: Torgashov, Nikita, et al.
Published: (2025)