PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khanal, Subash, Xing, Eric, Sastry, Srikumar, Dhakal, Aayush, Xiong, Zhexiao, Ahmad, Adeel, Jacobs, Nathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
von: Khanal, Subash, et al.
Veröffentlicht: (2025)
von: Khanal, Subash, et al.
Veröffentlicht: (2025)
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
von: Jogi, Yash, et al.
Veröffentlicht: (2025)
von: Jogi, Yash, et al.
Veröffentlicht: (2025)
TaxaBind: A Unified Embedding Space for Ecological Applications
von: Sastry, Srikumar, et al.
Veröffentlicht: (2024)
von: Sastry, Srikumar, et al.
Veröffentlicht: (2024)
RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings
von: Dhakal, Aayush, et al.
Veröffentlicht: (2025)
von: Dhakal, Aayush, et al.
Veröffentlicht: (2025)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Feature Selection via Graph Topology Inference for Soundscape Emotion Recognition
von: Rey, Samuel, et al.
Veröffentlicht: (2025)
von: Rey, Samuel, et al.
Veröffentlicht: (2025)
Autonomous Soundscape Augmentation with Multimodal Fusion of Visual and Participant-linked Inputs
von: Ooi, Kenneth, et al.
Veröffentlicht: (2023)
von: Ooi, Kenneth, et al.
Veröffentlicht: (2023)
Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation
von: Soltero, Kaspar, et al.
Veröffentlicht: (2025)
von: Soltero, Kaspar, et al.
Veröffentlicht: (2025)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
Zero-Shot Crate Digging: DJ Tool Retrieval Using Speech Activity, Music Structure And CLAP Embeddings
von: Orife, Iroro
Veröffentlicht: (2024)
von: Orife, Iroro
Veröffentlicht: (2024)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
ARAUS: A Large-Scale Dataset and Baseline Models of Affective Responses to Augmented Urban Soundscapes
von: Ooi, Kenneth, et al.
Veröffentlicht: (2022)
von: Ooi, Kenneth, et al.
Veröffentlicht: (2022)
Automating Urban Soundscape Enhancements with AI: In-situ Assessment of Quality and Restorativeness in Traffic-Exposed Residential Areas
von: Lam, Bhan, et al.
Veröffentlicht: (2024)
von: Lam, Bhan, et al.
Veröffentlicht: (2024)
Embedding-Space Diffusion for Zero-Shot Environmental Sound Classification
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
von: Chhibber, Manasi, et al.
Veröffentlicht: (2024)
Intelli-Z: Toward Intelligible Zero-Shot TTS
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
Fine-grained Soundscape Control for Augmented Hearing
von: Oh, Seunghyun, et al.
Veröffentlicht: (2026)
von: Oh, Seunghyun, et al.
Veröffentlicht: (2026)
Sound Tagging in Infant-centric Home Soundscapes
von: Khan, Mohammad Nur Hossain, et al.
Veröffentlicht: (2024)
von: Khan, Mohammad Nur Hossain, et al.
Veröffentlicht: (2024)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
GEOBIND: Binding Text, Image, and Audio through Satellite Images
von: Dhakal, Aayush, et al.
Veröffentlicht: (2024)
von: Dhakal, Aayush, et al.
Veröffentlicht: (2024)
Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
von: Dutta, Bikash, et al.
Veröffentlicht: (2025)
von: Dutta, Bikash, et al.
Veröffentlicht: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Few-Shot Bioacoustic Event Detection with Frame-Level Embedding Learning System
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
Generating Moving 3D Soundscapes with Latent Diffusion Models
von: Templin, Christian, et al.
Veröffentlicht: (2025)
von: Templin, Christian, et al.
Veröffentlicht: (2025)
Misophonia Trigger Sound Detection on Synthetic Soundscapes Using a Hybrid Model with a Frozen Pre-Trained CNN and a Time-Series Module
von: Sashida, Kurumi, et al.
Veröffentlicht: (2026)
von: Sashida, Kurumi, et al.
Veröffentlicht: (2026)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023)
von: Li, Junjie, et al.
Veröffentlicht: (2023)
Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
von: Li, Haowen, et al.
Veröffentlicht: (2026)
von: Li, Haowen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
von: Khanal, Subash, et al.
Veröffentlicht: (2025) -
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
von: Jogi, Yash, et al.
Veröffentlicht: (2025) -
TaxaBind: A Unified Embedding Space for Ecological Applications
von: Sastry, Srikumar, et al.
Veröffentlicht: (2024) -
RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings
von: Dhakal, Aayush, et al.
Veröffentlicht: (2025) -
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)