Gespeichert in:
| Hauptverfasser: | Bovbjerg, Holger Severin, Østergaard, Jan, Jensen, Jesper, Watanabe, Shinji, Tan, Zheng-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2508.20914 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023)
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
Noise-Robust Keyword Spotting through Self-supervised Pretraining
von: Mørk, Jacob, et al.
Veröffentlicht: (2024)
von: Mørk, Jacob, et al.
Veröffentlicht: (2024)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)
von: Li, Haowen, et al.
Veröffentlicht: (2025)
Audio-based Kinship Verification Using Age Domain Conversion
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
STAR: Speech-to-Audio Generation via Representation Learning
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
Symbolic Audio Classification via Modal Decision Tree Learning
von: Marzano, Enrico, et al.
Veröffentlicht: (2025)
von: Marzano, Enrico, et al.
Veröffentlicht: (2025)
ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
BAST: Binaural Audio Spectrogram Transformer for Binaural Sound Localization
von: Kuang, Sheng, et al.
Veröffentlicht: (2022)
von: Kuang, Sheng, et al.
Veröffentlicht: (2022)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
FakeSound: Deepfake General Audio Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Passive Underwater Acoustic Signal Separation based on Feature Decoupling Dual-path Network
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Deep low-latency joint speech transmission and enhancement over a gaussian channel
von: Bokaei, Mohammad, et al.
Veröffentlicht: (2024)
von: Bokaei, Mohammad, et al.
Veröffentlicht: (2024)
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
SARA: Stress Test Reasoning in Audio Deepfake Detection
von: Nguyen, Binh, et al.
Veröffentlicht: (2026)
von: Nguyen, Binh, et al.
Veröffentlicht: (2026)
Joint Minimum Processing Beamforming and Near-end Listening Enhancement
von: Fuglsig, Andreas J., et al.
Veröffentlicht: (2023)
von: Fuglsig, Andreas J., et al.
Veröffentlicht: (2023)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
von: Dhiman, Jai
Veröffentlicht: (2026)
von: Dhiman, Jai
Veröffentlicht: (2026)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
von: Khushiyant, et al.
Veröffentlicht: (2026)
von: Khushiyant, et al.
Veröffentlicht: (2026)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Investigating the Design Space of Diffusion Models for Speech Enhancement
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed Tokenization
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2025)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Online Single-Channel Audio-Based Sound Speed Estimation for Robust Multi-Channel Audio Control
von: Fuglsig, Andreas Jonas, et al.
Veröffentlicht: (2026)
von: Fuglsig, Andreas Jonas, et al.
Veröffentlicht: (2026)
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
von: Aristorenas, Aris J.
Veröffentlicht: (2024)
von: Aristorenas, Aris J.
Veröffentlicht: (2024)
TuneGenie: Reasoning-based LLM agents for preferential music generation
von: Pandey, Amitesh, et al.
Veröffentlicht: (2025)
von: Pandey, Amitesh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023) -
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025) -
Noise-Robust Keyword Spotting through Self-supervised Pretraining
von: Mørk, Jacob, et al.
Veröffentlicht: (2024) -
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026) -
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)