Establishing degrees of closeness between audio recordings along different dimensions using large-scale cross-lingual models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fily, Maxime, Wisniewski, Guillaume, Guillaume, Severine, Adda, Gilles, Michaud, Alexis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
von: Wisniewski, Guillaume, et al.
Veröffentlicht: (2025)
von: Wisniewski, Guillaume, et al.
Veröffentlicht: (2025)
ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
von: Niu, Yadong, et al.
Veröffentlicht: (2026)
von: Niu, Yadong, et al.
Veröffentlicht: (2026)
Combining audio control and style transfer using latent diffusion
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
Scaling up masked audio encoder learning for general audio classification
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
Joint Multi-scale Cross-lingual Speaking Style Transfer with Bidirectional Attention Mechanism for Automatic Dubbing
von: Li, Jingbei, et al.
Veröffentlicht: (2023)
von: Li, Jingbei, et al.
Veröffentlicht: (2023)
Towards audio language modeling -- an overview
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Are audio DeepFake detection models polyglots?
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
Tweaking autoregressive methods for inpainting of gaps in audio signals
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
MBCodec:Thorough disentangle for high-fidelity audio compression
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
Real-time implementation of vibrato transfer as an audio effect
von: Hyrkas, Jeremy
Veröffentlicht: (2025)
von: Hyrkas, Jeremy
Veröffentlicht: (2025)
Robustness assessment of large audio language models in multiple-choice evaluation
von: López, Fernando, et al.
Veröffentlicht: (2025)
von: López, Fernando, et al.
Veröffentlicht: (2025)
ADIFF: Explaining audio difference using natural language
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
FxSearcher: gradient-free text-driven audio transformation
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
Regularized autoregressive modeling and its application to audio signal reconstruction
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
EDTC: enhance depth of text comprehension in automated audio captioning
von: Tan, Liwen, et al.
Veröffentlicht: (2024)
von: Tan, Liwen, et al.
Veröffentlicht: (2024)
AxLSTMs: learning self-supervised audio representations with xLSTMs
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
STASE: A spatialized text-to-audio synthesis engine for music generation
von: Chi, Tutti, et al.
Veröffentlicht: (2025)
von: Chi, Tutti, et al.
Veröffentlicht: (2025)
Deep learning based spatial aliasing reduction in beamforming for audio capture
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
von: Das, Sneha, et al.
Veröffentlicht: (2020)
von: Das, Sneha, et al.
Veröffentlicht: (2020)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
DashengTokenizer: One layer is enough for unified audio understanding and generation
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
ICGAN: An implicit conditioning method for interpretable feature control of neural audio synthesis
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
A robust audio deepfake detection system via multi-view feature
von: Yang, Yujie, et al.
Veröffentlicht: (2024)
von: Yang, Yujie, et al.
Veröffentlicht: (2024)
Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
von: Riley, Xavier, et al.
Veröffentlicht: (2024)
von: Riley, Xavier, et al.
Veröffentlicht: (2024)
Exploring trends in audio mixes and masters: Insights from a dataset analysis
von: Mourgela, Angeliki, et al.
Veröffentlicht: (2024)
von: Mourgela, Angeliki, et al.
Veröffentlicht: (2024)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes
von: Hsu, Yicheng, et al.
Veröffentlicht: (2024)
von: Hsu, Yicheng, et al.
Veröffentlicht: (2024)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
audio2chart: End to End Audio Transcription into playable Guitar Hero charts
von: Tripodi, Riccardo
Veröffentlicht: (2025)
von: Tripodi, Riccardo
Veröffentlicht: (2025)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension
von: Gupta, Kishan, et al.
Veröffentlicht: (2025)
von: Gupta, Kishan, et al.
Veröffentlicht: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
von: Han, Runduo, et al.
Veröffentlicht: (2024)
von: Han, Runduo, et al.
Veröffentlicht: (2024)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
von: Wisniewski, Guillaume, et al.
Veröffentlicht: (2025) -
ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
von: Niu, Yadong, et al.
Veröffentlicht: (2026) -
Combining audio control and style transfer using latent diffusion
von: Demerlé, Nils, et al.
Veröffentlicht: (2024) -
Scaling up masked audio encoder learning for general audio classification
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024) -
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)