Multiple output samples per input in a single-output Gaussian process
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wong, Jeremy H. M., Zhang, Huayun, Chen, Nancy F. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Design framework for spherical microphone and loudspeaker arrays in a multiple-input multiple-output system
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
DistriBlock: Identifying adversarial audio samples by leveraging characteristics of the output distribution
von: Pizarro, Matías, et al.
Veröffentlicht: (2023)
von: Pizarro, Matías, et al.
Veröffentlicht: (2023)
Theory and investigation of acoustic multiple-input multiple-output systems based on spherical arrays in a room
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio
von: AbdulKader, Ahmad, et al.
Veröffentlicht: (2017)
von: AbdulKader, Ahmad, et al.
Veröffentlicht: (2017)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
Audio-to-Score Conversion Model Based on Whisper methodology
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
von: Falai, Alessio, et al.
Veröffentlicht: (2025)
von: Falai, Alessio, et al.
Veröffentlicht: (2025)
Generative Pre-training for Speech with Flow Matching
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
The State Of TTS: A Case Study with Human Fooling Rates
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
Evaluation of Google's Voice Recognition and Sentence Classification for Health Care Applications
von: Uddin, Majbah, et al.
Veröffentlicht: (2024)
von: Uddin, Majbah, et al.
Veröffentlicht: (2024)
Multiple Choice Learning for Efficient Speech Separation with Many Speakers
von: Perera, David, et al.
Veröffentlicht: (2024)
von: Perera, David, et al.
Veröffentlicht: (2024)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2025)
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2025)
CleanCTG: A Deep Learning Model for Multi-Artefact Detection and Reconstruction in Cardiotocography
von: Wong, Sheng, et al.
Veröffentlicht: (2025)
von: Wong, Sheng, et al.
Veröffentlicht: (2025)
An Empirical Recipe for Universal Phone Recognition
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
von: Tankasala, Srinath, et al.
Veröffentlicht: (2023)
von: Tankasala, Srinath, et al.
Veröffentlicht: (2023)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
von: Anand, Srija, et al.
Veröffentlicht: (2024)
von: Anand, Srija, et al.
Veröffentlicht: (2024)
Audio Simulation for Sound Source Localization in Virtual Evironment
von: Di Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Di Yuan, Yi, et al.
Veröffentlicht: (2024)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
von: Lei, Zhihong, et al.
Veröffentlicht: (2024)
von: Lei, Zhihong, et al.
Veröffentlicht: (2024)
HiSSNet: Sound Event Detection and Speaker Identification via Hierarchical Prototypical Networks for Low-Resource Headphones
von: Shashaank, N, et al.
Veröffentlicht: (2023)
von: Shashaank, N, et al.
Veröffentlicht: (2023)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2023)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Design framework for spherical microphone and loudspeaker arrays in a multiple-input multiple-output system
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024) -
DistriBlock: Identifying adversarial audio samples by leveraging characteristics of the output distribution
von: Pizarro, Matías, et al.
Veröffentlicht: (2023) -
Theory and investigation of acoustic multiple-input multiple-output systems based on spherical arrays in a room
von: Morgenstern, Hai, et al.
Veröffentlicht: (2024) -
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022) -
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)