DLIOS: An LLM-Augmented Real-Time Multi-Modal Interactive Enhancement Overlay System for Douyin Live Streaming
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Shuide, Seok, Sungil, Ku, Beier, Li, Richee, He, Yubin, Qu, Bowen, Yang, Yang, Su, Ping, Jiao, Can |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
by: Yang, Gene-Ping, et al.
Published: (2025)
by: Yang, Gene-Ping, et al.
Published: (2025)
Data Augmentation for Pathological Speech Enhancement
by: Hou, Mingchi, et al.
Published: (2026)
by: Hou, Mingchi, et al.
Published: (2026)
FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement
by: Ahn, Sunghwan, et al.
Published: (2025)
by: Ahn, Sunghwan, et al.
Published: (2025)
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
by: Ho, Chun-Wei, et al.
Published: (2025)
by: Ho, Chun-Wei, et al.
Published: (2025)
UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
by: Rong, Xiaobin, et al.
Published: (2025)
by: Rong, Xiaobin, et al.
Published: (2025)
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
by: Yang, Bing, et al.
Published: (2024)
by: Yang, Bing, et al.
Published: (2024)
Test-Time Adaptation For Speech Enhancement Via Mask Polarization
by: Raichle, Tobias, et al.
Published: (2026)
by: Raichle, Tobias, et al.
Published: (2026)
Dynamically Slimmable Speech Enhancement Network with Metric-Guided Training
by: Zhao, Haixin, et al.
Published: (2025)
by: Zhao, Haixin, et al.
Published: (2025)
SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
by: Jiang, Shaohan, et al.
Published: (2025)
by: Jiang, Shaohan, et al.
Published: (2025)
Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet
by: Hao, Xiang, et al.
Published: (2024)
by: Hao, Xiang, et al.
Published: (2024)
Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation
by: Raichle, Tobias, et al.
Published: (2025)
by: Raichle, Tobias, et al.
Published: (2025)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
by: Yang, Mu, et al.
Published: (2024)
by: Yang, Mu, et al.
Published: (2024)
SE Territory: Monaural Speech Enhancement Meets the Fixed Virtual Perceptual Space Mapping
by: Xu, Xinmeng, et al.
Published: (2023)
by: Xu, Xinmeng, et al.
Published: (2023)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
by: Ku, Pin-Jui, et al.
Published: (2024)
by: Ku, Pin-Jui, et al.
Published: (2024)
Robust Nasality Representation Learning for Cleft Palate-Related Velopharyngeal Dysfunction Screening in Real-World Settings
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
by: Quan, Changsheng, et al.
Published: (2024)
by: Quan, Changsheng, et al.
Published: (2024)
Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure
by: Togami, Masahito, et al.
Published: (2024)
by: Togami, Masahito, et al.
Published: (2024)
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
by: Lu, Ye-Xin, et al.
Published: (2023)
by: Lu, Ye-Xin, et al.
Published: (2023)
Hybrid Real- And Complex-Valued Neural Network Concept For Low-Complexity Phase-Aware Speech Enhancement
by: Fiorio, Luan Vinícius, et al.
Published: (2025)
by: Fiorio, Luan Vinícius, et al.
Published: (2025)
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
MeanSE: Efficient Generative Speech Enhancement with Mean Flows
by: Wang, Jiahe, et al.
Published: (2025)
by: Wang, Jiahe, et al.
Published: (2025)
Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement
by: Wang, Tingting, et al.
Published: (2024)
by: Wang, Tingting, et al.
Published: (2024)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
by: Yang, Yujie, et al.
Published: (2025)
by: Yang, Yujie, et al.
Published: (2025)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
by: La Quatra, Moreno, et al.
Published: (2024)
by: La Quatra, Moreno, et al.
Published: (2024)
An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding
by: Su, Tianhui, et al.
Published: (2026)
by: Su, Tianhui, et al.
Published: (2026)
Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
by: Yang, Mu, et al.
Published: (2025)
by: Yang, Mu, et al.
Published: (2025)
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
by: Yang, Hsiang-Cheng, et al.
Published: (2026)
by: Yang, Hsiang-Cheng, et al.
Published: (2026)
High-Density MIMO Localization Using a 32x64 Ultrasonic Transducer-Microphone Array with Real-Time Data Streaming
by: Baeyens, Rens, et al.
Published: (2025)
by: Baeyens, Rens, et al.
Published: (2025)
Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
by: Kirton-Wingate, Jasper, et al.
Published: (2024)
by: Kirton-Wingate, Jasper, et al.
Published: (2024)
Speech Loudness in Broadcasting and Streaming
by: Torcoli, Matteo, et al.
Published: (2024)
by: Torcoli, Matteo, et al.
Published: (2024)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Unsupervised Face-Masked Speech Enhancement Using Generative Adversarial Networks With Human-in-the-Loop Assessment Metrics
by: Wang, Syu-Siang, et al.
Published: (2024)
by: Wang, Syu-Siang, et al.
Published: (2024)
nlm: Real-Time Non-linear Modal Synthesis in Max
by: Diaz, Rodrigo, et al.
Published: (2026)
by: Diaz, Rodrigo, et al.
Published: (2026)
T-Mimi: A Transformer-based Mimi Decoder for Real-Time On-Phone TTS
by: Wu, Haibin, et al.
Published: (2026)
by: Wu, Haibin, et al.
Published: (2026)
StreamVC: Real-Time Low-Latency Voice Conversion
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
Chunkwise Aligners for Streaming Speech Recognition
by: Teo, Wen Shen, et al.
Published: (2026)
by: Teo, Wen Shen, et al.
Published: (2026)
Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement
by: Li, Yi, et al.
Published: (2024)
by: Li, Yi, et al.
Published: (2024)
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
by: Jiang, Xiao-Hang, et al.
Published: (2026)
by: Jiang, Xiao-Hang, et al.
Published: (2026)
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
by: Korse, Srikanth, et al.
Published: (2025)
by: Korse, Srikanth, et al.
Published: (2025)
Similar Items
-
Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
by: Yang, Gene-Ping, et al.
Published: (2025) -
Data Augmentation for Pathological Speech Enhancement
by: Hou, Mingchi, et al.
Published: (2026) -
FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement
by: Ahn, Sunghwan, et al.
Published: (2025) -
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
by: Ho, Chun-Wei, et al.
Published: (2025) -
UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
by: Rong, Xiaobin, et al.
Published: (2025)