STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wisnu, Dyah A. M. G., Zezario, Ryandhimas E., Rini, Stefano, Li, Fo-Rui, Peng, Yan-Tsung, Wang, Hsin-Min, Tsao, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HAAQI-Net: A Non-intrusive Neural Music Audio Quality Assessment Model for Hearing Aids
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2024)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2024)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution
von: Abreu, Wallace, et al.
Veröffentlicht: (2026)
von: Abreu, Wallace, et al.
Veröffentlicht: (2026)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
von: Zezario, Ryandhimas E.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E.
Veröffentlicht: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
A Study on Incorporating Whisper for Robust Speech Assessment
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
A Study on Speech Assessment with Visual Cues
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2021)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2021)
Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
von: Yuan, Hui-Guan, et al.
Veröffentlicht: (2025)
von: Yuan, Hui-Guan, et al.
Veröffentlicht: (2025)
NeuroAMP: A Novel End-to-end General Purpose Deep Neural Amplifier for Personalized Hearing Aids
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
StressTest: Can YOUR Speech LM Handle the Stress?
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
von: Nakazawa, Kazushi
Veröffentlicht: (2026)
von: Nakazawa, Kazushi
Veröffentlicht: (2026)
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
OpusLM: A Family of Open Unified Speech Language Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
DroFiT: A Lightweight Band-fused Frequency Attention Toward Real-time UAV Speech Enhancement
von: Lee, Jeongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jeongmin, et al.
Veröffentlicht: (2025)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Bridging the gap between training and inference in LM-based TTS models
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial Networks
von: Dmitrieva, Ekaterina, et al.
Veröffentlicht: (2025)
von: Dmitrieva, Ekaterina, et al.
Veröffentlicht: (2025)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
von: Yin, Chun, et al.
Veröffentlicht: (2024)
von: Yin, Chun, et al.
Veröffentlicht: (2024)
SonicVisionLM: Playing Sound with Vision Language Models
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
Spirit LM: Interleaved Spoken and Written Language Model
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HAAQI-Net: A Non-intrusive Neural Music Audio Quality Assessment Model for Hearing Aids
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2024) -
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025) -
A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025) -
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025) -
FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution
von: Abreu, Wallace, et al.
Veröffentlicht: (2026)