Saved in:
| Main Authors: | Chen, Tuochao, Wang, Qirui, Wu, Bohan, Itani, Malek, Eskimez, Sefik Emre, Yoshioka, Takuya, Gollakota, Shyamnath |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.11277 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge boosting during low-latency inference
by: Srinivas, Vidya, et al.
Published: (2024)
by: Srinivas, Vidya, et al.
Published: (2024)
Neural Speech Extraction with Human Feedback
by: Itani, Malek, et al.
Published: (2025)
by: Itani, Malek, et al.
Published: (2025)
Look Once to Hear: Target Speech Hearing with Noisy Examples
by: Veluri, Bandhav, et al.
Published: (2024)
by: Veluri, Bandhav, et al.
Published: (2024)
TF-MLPNet: Tiny Real-Time Neural Speech Separation
by: Itani, Malek, et al.
Published: (2025)
by: Itani, Malek, et al.
Published: (2025)
Proactive Hearing Assistants that Isolate Egocentric Conversations
by: Hu, Guilin, et al.
Published: (2025)
by: Hu, Guilin, et al.
Published: (2025)
Wireless Hearables With Programmable Speech AI Accelerators
by: Itani, Malek, et al.
Published: (2025)
by: Itani, Malek, et al.
Published: (2025)
Fine-grained Soundscape Control for Augmented Hearing
by: Oh, Seunghyun, et al.
Published: (2026)
by: Oh, Seunghyun, et al.
Published: (2026)
TS3-Codec: Transformer-Based Simple Streaming Single Codec
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
Spatial Speech Translation: Translating Across Space With Binaural Hearables
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
LLAMAPIE: Proactive In-Ear Conversation Assistants
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
by: Wang, Xiaofei, et al.
Published: (2023)
by: Wang, Xiaofei, et al.
Published: (2023)
Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
by: Eskimez, Sefik Emre, et al.
Published: (2024)
by: Eskimez, Sefik Emre, et al.
Published: (2024)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
by: Veluri, Bandhav, et al.
Published: (2024)
by: Veluri, Bandhav, et al.
Published: (2024)
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
by: Wang, Xiaofei, et al.
Published: (2024)
by: Wang, Xiaofei, et al.
Published: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
by: Eskimez, Sefik Emre, et al.
Published: (2024)
by: Eskimez, Sefik Emre, et al.
Published: (2024)
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
by: Wang, Dongmei, et al.
Published: (2023)
by: Wang, Dongmei, et al.
Published: (2023)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
Analysis and Extension of Noisy-target Training for Unsupervised Target Signal Enhancement
by: Fujimura, Takuya, et al.
Published: (2025)
by: Fujimura, Takuya, et al.
Published: (2025)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
by: Lin, Guan-Ting, et al.
Published: (2025)
by: Lin, Guan-Ting, et al.
Published: (2025)
Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
by: Kanda, Naoyuki, et al.
Published: (2024)
by: Kanda, Naoyuki, et al.
Published: (2024)
DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
by: Hasumi, Takuya, et al.
Published: (2025)
by: Hasumi, Takuya, et al.
Published: (2025)
Turn-taking annotation for quantitative and qualitative analyses of conversation
by: Kelterer, Anneliese, et al.
Published: (2025)
by: Kelterer, Anneliese, et al.
Published: (2025)
Comparison of Classification Algorithms for COVID19 Detection using Cough Acoustic Signals
by: Erdoğan, Yunus Emre, et al.
Published: (2022)
by: Erdoğan, Yunus Emre, et al.
Published: (2022)
Transformer-based End-to-End Control Filter Generation for Active Noise Control
by: Yang, Ziyi, et al.
Published: (2026)
by: Yang, Ziyi, et al.
Published: (2026)
Neural Network-Based Time-Frequency-Bin-Wise Linear Combination of Beamformers for Underdetermined Target Source Extraction
by: Chen, Changda, et al.
Published: (2026)
by: Chen, Changda, et al.
Published: (2026)
MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
by: Chen, Hongzhao, et al.
Published: (2025)
by: Chen, Hongzhao, et al.
Published: (2025)
Investigating self-supervised features for expressive, multilingual voice conversion
by: Martín-Cortinas, Álvaro, et al.
Published: (2025)
by: Martín-Cortinas, Álvaro, et al.
Published: (2025)
A Knowledge-Driven Approach to Target Speech Extraction in the Presence of Background Sound Effects for Cinematic Audio Source Separation (CASS)
by: Ho, Chun-wei, et al.
Published: (2026)
by: Ho, Chun-wei, et al.
Published: (2026)
Why does music source separation benefit from cacophony?
by: Jeon, Chang-Bin, et al.
Published: (2024)
by: Jeon, Chang-Bin, et al.
Published: (2024)
Real-Time and Accurate: Zero-shot High-Fidelity Singing Voice Conversion with Multi-Condition Flow Synthesis
by: Li, Hui, et al.
Published: (2024)
by: Li, Hui, et al.
Published: (2024)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
by: Huang, Ziling, et al.
Published: (2025)
by: Huang, Ziling, et al.
Published: (2025)
ARTT: Augmented Reverberant-Target Training for Unsupervised Monaural Speech Dereverberation
by: Song, Siqi, et al.
Published: (2026)
by: Song, Siqi, et al.
Published: (2026)
BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
by: Wang, Haoran, et al.
Published: (2025)
by: Wang, Haoran, et al.
Published: (2025)
SCNet: Sparse Compression Network for Music Source Separation
by: Tong, Weinan, et al.
Published: (2024)
by: Tong, Weinan, et al.
Published: (2024)
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
by: Wu, Ke, et al.
Published: (2026)
by: Wu, Ke, et al.
Published: (2026)
Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
by: Lin, Guan-Ting, et al.
Published: (2025)
by: Lin, Guan-Ting, et al.
Published: (2025)
Multiple Speaker Separation from Noisy Sources in Reverberant Rooms using Relative Transfer Matrix
by: Manamperi, Wageesha N., et al.
Published: (2025)
by: Manamperi, Wageesha N., et al.
Published: (2025)
Leveraging Sound Source Trajectories for Universal Sound Separation
by: Wu, Donghang, et al.
Published: (2024)
by: Wu, Donghang, et al.
Published: (2024)
Low algorithmic delay implementation of convolutional beamformer for online joint source separation and dereverberation
by: Mo, Kaien, et al.
Published: (2024)
by: Mo, Kaien, et al.
Published: (2024)
Similar Items
-
Knowledge boosting during low-latency inference
by: Srinivas, Vidya, et al.
Published: (2024) -
Neural Speech Extraction with Human Feedback
by: Itani, Malek, et al.
Published: (2025) -
Look Once to Hear: Target Speech Hearing with Noisy Examples
by: Veluri, Bandhav, et al.
Published: (2024) -
TF-MLPNet: Tiny Real-Time Neural Speech Separation
by: Itani, Malek, et al.
Published: (2025) -
Proactive Hearing Assistants that Isolate Egocentric Conversations
by: Hu, Guilin, et al.
Published: (2025)