Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Jaeyeon, Hwang, Injune, Lee, Kyogu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
di: Hwang, Injune, et al.
Pubblicazione: (2024)
di: Hwang, Injune, et al.
Pubblicazione: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
di: Han, Seungu, et al.
Pubblicazione: (2026)
di: Han, Seungu, et al.
Pubblicazione: (2026)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
ViSAGe: Video-to-Spatial Audio Generation
di: Kim, Jaeyeon, et al.
Pubblicazione: (2025)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2025)
Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training
di: Joung, Haesun, et al.
Pubblicazione: (2024)
di: Joung, Haesun, et al.
Pubblicazione: (2024)
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
Hear Your Face: Face-based voice conversion with F0 estimation
di: Lee, Jaejun, et al.
Pubblicazione: (2024)
di: Lee, Jaejun, et al.
Pubblicazione: (2024)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
di: Kim, Jaeyeon, et al.
Pubblicazione: (2025)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2025)
Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
di: Rhyu, Seungyeon, et al.
Pubblicazione: (2024)
di: Rhyu, Seungyeon, et al.
Pubblicazione: (2024)
Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
di: Jang, Jaehyuk, et al.
Pubblicazione: (2026)
di: Jang, Jaehyuk, et al.
Pubblicazione: (2026)
Generalizable Audio Spoofing Detection using Non-Semantic Representations
di: Das, Arnab, et al.
Pubblicazione: (2025)
di: Das, Arnab, et al.
Pubblicazione: (2025)
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
di: Chen, Shunian, et al.
Pubblicazione: (2025)
di: Chen, Shunian, et al.
Pubblicazione: (2025)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
di: Erol, Mehmet Hamza, et al.
Pubblicazione: (2024)
di: Erol, Mehmet Hamza, et al.
Pubblicazione: (2024)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
di: Baser, Oguzhan, et al.
Pubblicazione: (2025)
di: Baser, Oguzhan, et al.
Pubblicazione: (2025)
Audio Signal Processing Using Time Domain Mel-Frequency Wavelet Coefficient
di: Sebastian, Rinku, et al.
Pubblicazione: (2025)
di: Sebastian, Rinku, et al.
Pubblicazione: (2025)
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
di: Han, Bing, et al.
Pubblicazione: (2026)
di: Han, Bing, et al.
Pubblicazione: (2026)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
di: Ma, Yi, et al.
Pubblicazione: (2025)
di: Ma, Yi, et al.
Pubblicazione: (2025)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
di: Haque, Kazi Nazmul, et al.
Pubblicazione: (2024)
di: Haque, Kazi Nazmul, et al.
Pubblicazione: (2024)
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
di: Oh, Yoori, et al.
Pubblicazione: (2024)
di: Oh, Yoori, et al.
Pubblicazione: (2024)
DOSE : Drum One-Shot Extraction from Music Mixture
di: Hwang, Suntae, et al.
Pubblicazione: (2025)
di: Hwang, Suntae, et al.
Pubblicazione: (2025)
BrewCLIP: A Bifurcated Representation Learning Framework for Audio-Visual Retrieval
di: Lu, Zhenyu, et al.
Pubblicazione: (2024)
di: Lu, Zhenyu, et al.
Pubblicazione: (2024)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
di: Luong, Justin, et al.
Pubblicazione: (2025)
di: Luong, Justin, et al.
Pubblicazione: (2025)
Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
di: Guo, Zixun, et al.
Pubblicazione: (2025)
di: Guo, Zixun, et al.
Pubblicazione: (2025)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
Automated Classification of Phonetic Segments in Child Speech Using Raw Ultrasound Imaging
di: Ani, Saja Al, et al.
Pubblicazione: (2024)
di: Ani, Saja Al, et al.
Pubblicazione: (2024)
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
di: Mahmud, Tanvir, et al.
Pubblicazione: (2024)
di: Mahmud, Tanvir, et al.
Pubblicazione: (2024)
Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
di: Deng, Qixin, et al.
Pubblicazione: (2025)
di: Deng, Qixin, et al.
Pubblicazione: (2025)
Do Captioning Metrics Reflect Music Semantic Alignment?
di: Lee, Jinwoo, et al.
Pubblicazione: (2024)
di: Lee, Jinwoo, et al.
Pubblicazione: (2024)
Contextual Biasing to Improve Domain-specific Custom Vocabulary Audio Transcription without Explicit Fine-Tuning of Whisper Model
di: Lall, Vishakha, et al.
Pubblicazione: (2024)
di: Lall, Vishakha, et al.
Pubblicazione: (2024)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
di: Kuznetsova, Anastasia, et al.
Pubblicazione: (2025)
di: Kuznetsova, Anastasia, et al.
Pubblicazione: (2025)
Token Pruning in Audio Transformers: Optimizing Performance and Decoding Patch Importance
di: Lee, Taehan, et al.
Pubblicazione: (2025)
di: Lee, Taehan, et al.
Pubblicazione: (2025)
DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2025)
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2025)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
di: Zhang, Yaoyun, et al.
Pubblicazione: (2024)
di: Zhang, Yaoyun, et al.
Pubblicazione: (2024)
Wavespace: A Highly Explorable Wavetable Generator
di: Lee, Hazounne, et al.
Pubblicazione: (2024)
di: Lee, Hazounne, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
di: Hwang, Injune, et al.
Pubblicazione: (2024) -
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
di: Han, Seungu, et al.
Pubblicazione: (2026) -
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024) -
ViSAGe: Video-to-Spatial Audio Generation
di: Kim, Jaeyeon, et al.
Pubblicazione: (2025) -
Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training
di: Joung, Haesun, et al.
Pubblicazione: (2024)