The language of sound search: Examining User Queries in Audio Search Engines
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Weck, Benno, Font, Frederic |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
von: Weck, Benno, et al.
Veröffentlicht: (2023)
von: Weck, Benno, et al.
Veröffentlicht: (2023)
Navigating Speech Recording Collections with AI-Generated Illustrations
von: Håland, Sirina, et al.
Veröffentlicht: (2025)
von: Håland, Sirina, et al.
Veröffentlicht: (2025)
Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
von: Garcia, Nelly, et al.
Veröffentlicht: (2026)
von: Garcia, Nelly, et al.
Veröffentlicht: (2026)
Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
von: Baranes, Marion, et al.
Veröffentlicht: (2026)
von: Baranes, Marion, et al.
Veröffentlicht: (2026)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
A conversational gesture synthesis system based on emotions and semantics
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
LLAMAPIE: Proactive In-Ear Conversation Assistants
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
Literary and Colloquial Tamil Dialect Identification
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
von: Liu, Xing Yi, et al.
Veröffentlicht: (2023)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2023)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
Exploring Situated Stabilities of a Rhythm Generation System through Variational Cross-Examination
von: Kotowski, Błażej, et al.
Veröffentlicht: (2025)
von: Kotowski, Błażej, et al.
Veröffentlicht: (2025)
Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Technical Report on classification of literature related to children speech disorder
von: Wang, Ziang, et al.
Veröffentlicht: (2025)
von: Wang, Ziang, et al.
Veröffentlicht: (2025)
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2024)
von: Vasilakis, Yannis, et al.
Veröffentlicht: (2024)
More than words: Advancements and challenges in speech recognition for singing
von: Kruspe, Anna
Veröffentlicht: (2024)
von: Kruspe, Anna
Veröffentlicht: (2024)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
Towards Reliable Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
The effect of self-motion and room familiarity on sound source localization in virtual environments
von: Isserstedt, Niklas, et al.
Veröffentlicht: (2024)
von: Isserstedt, Niklas, et al.
Veröffentlicht: (2024)
ExSampling: a system for the real-time ensemble performance of field-recorded environmental sounds
von: Kobayashi, Atsuya, et al.
Veröffentlicht: (2020)
von: Kobayashi, Atsuya, et al.
Veröffentlicht: (2020)
Hidden bawls, whispers, and yelps: can text be made to sound more than just its words?
von: Pataca, Caluã de Lacerda, et al.
Veröffentlicht: (2022)
von: Pataca, Caluã de Lacerda, et al.
Veröffentlicht: (2022)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
von: Jaiswal, Pranay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Pranay, et al.
Veröffentlicht: (2024)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
Improving AI-generated music with user-guided training
von: Singh, Vishwa Mohan, et al.
Veröffentlicht: (2025)
von: Singh, Vishwa Mohan, et al.
Veröffentlicht: (2025)
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition
von: Kim, Hyunju, et al.
Veröffentlicht: (2024)
von: Kim, Hyunju, et al.
Veröffentlicht: (2024)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
von: Doukhan, David, et al.
Veröffentlicht: (2024)
von: Doukhan, David, et al.
Veröffentlicht: (2024)
Seeing Beyond Sound: Visualization and Abstraction in Audio Data Representation
von: Blum'e, Ashlae
Veröffentlicht: (2025)
von: Blum'e, Ashlae
Veröffentlicht: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing System
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2025)
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2025)
Learning Relationships Between Separate Audio Tracks for Creative Applications
von: Bujard, Balthazar, et al.
Veröffentlicht: (2025)
von: Bujard, Balthazar, et al.
Veröffentlicht: (2025)
A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
von: Zheng, Shuoyang, et al.
Veröffentlicht: (2024)
von: Zheng, Shuoyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
von: Weck, Benno, et al.
Veröffentlicht: (2023) -
Navigating Speech Recording Collections with AI-Generated Illustrations
von: Håland, Sirina, et al.
Veröffentlicht: (2025) -
Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
von: Garcia, Nelly, et al.
Veröffentlicht: (2026) -
Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
von: Baranes, Marion, et al.
Veröffentlicht: (2026) -
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)