VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
Fuente:
arXiv
Saved in:
| Main Authors: | Basha, Maris, Zai, Anja, Stoll, Sabine, Hahnloser, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real-World En Call Center Transcripts Dataset with PII Redaction
by: Dao, Ha, et al.
Published: (2025)
by: Dao, Ha, et al.
Published: (2025)
The Concatenator: A Bayesian Approach To Real Time Concatenative Musaicing
by: Tralie, Christopher, et al.
Published: (2024)
by: Tralie, Christopher, et al.
Published: (2024)
Deep Interest Mining for Intent-Enriched Semantic IDs in Multimodal Generative Recommendation
by: Zeng, Yangchen, et al.
Published: (2026)
by: Zeng, Yangchen, et al.
Published: (2026)
Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling
by: Beskorovainyi, Vladimir
Published: (2026)
by: Beskorovainyi, Vladimir
Published: (2026)
Fine-tuning Pre-trained Audio Models for COVID-19 Detection: A Technical Report
by: de Brito, Daniel Oliveira, et al.
Published: (2025)
by: de Brito, Daniel Oliveira, et al.
Published: (2025)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
by: Chien, Sheng-You, et al.
Published: (2026)
by: Chien, Sheng-You, et al.
Published: (2026)
InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking
by: Seetharaman, Rahul, et al.
Published: (2025)
by: Seetharaman, Rahul, et al.
Published: (2025)
Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation
by: Cheng, Shihan, et al.
Published: (2025)
by: Cheng, Shihan, et al.
Published: (2025)
Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback
by: Jahanbin, Peyman
Published: (2025)
by: Jahanbin, Peyman
Published: (2025)
INESC-ID @ eRisk 2025: Exploring Fine-Tuned, Similarity-Based, and Prompt-Based Approaches to Depression Symptom Identification
by: Nunes, Diogo A. P., et al.
Published: (2025)
by: Nunes, Diogo A. P., et al.
Published: (2025)
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
by: Ou, Yiwei, et al.
Published: (2026)
by: Ou, Yiwei, et al.
Published: (2026)
Spherical Hermite Maps
by: Abouagour, Mohamed, et al.
Published: (2026)
by: Abouagour, Mohamed, et al.
Published: (2026)
RARR : Robust Real-World Activity Recognition with Vibration by Scavenging Near-Surface Audio Online
by: Lee, Dong Yoon, et al.
Published: (2025)
by: Lee, Dong Yoon, et al.
Published: (2025)
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
by: Husain, Jaavid Aktar, et al.
Published: (2026)
by: Husain, Jaavid Aktar, et al.
Published: (2026)
Real-time Level-of-Detail Strand-based Hair Rendering
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models
by: Tralie, Christopher J., et al.
Published: (2024)
by: Tralie, Christopher J., et al.
Published: (2024)
AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
by: Zhang, Pengfei, et al.
Published: (2026)
by: Zhang, Pengfei, et al.
Published: (2026)
A study on audio synchronous steganography detection and distributed guide inference model based on sliding spectral features and intelligent inference drive
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Improving Angular Speed Uniformity by Piecewise Radical Reparameterization
by: Hong, Hoon, et al.
Published: (2024)
by: Hong, Hoon, et al.
Published: (2024)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
A ripple in time: a discontinuity in American history
by: Kolpakov, Alexander, et al.
Published: (2023)
by: Kolpakov, Alexander, et al.
Published: (2023)
A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
by: Károly, Artúr I., et al.
Published: (2025)
by: Károly, Artúr I., et al.
Published: (2025)
MelodyVis: Visual Analytics for Melodic Patterns in Sheet Music
by: Miller, Matthias, et al.
Published: (2024)
by: Miller, Matthias, et al.
Published: (2024)
Robust Average Networks for Monte Carlo Denoising
by: Kalojanov, Javor, et al.
Published: (2023)
by: Kalojanov, Javor, et al.
Published: (2023)
Polar Stroking: New Theory and Methods for Stroking Paths
by: Kilgard, Mark J.
Published: (2020)
by: Kilgard, Mark J.
Published: (2020)
Specular Polynomials
by: Fan, Zhimin, et al.
Published: (2024)
by: Fan, Zhimin, et al.
Published: (2024)
InverseVis: Revealing the Hidden with Curved Sphere Tracing
by: Lawonn, Kai, et al.
Published: (2024)
by: Lawonn, Kai, et al.
Published: (2024)
Interactive Visualization on Large High-Resolution Displays: A Survey
by: Belkacem, Ilyasse, et al.
Published: (2022)
by: Belkacem, Ilyasse, et al.
Published: (2022)
Apictorial Jigsaw Puzzle Reconstruction Based on Curve Matching via a Corotational Beam Spline
by: Orynyak, Igor, et al.
Published: (2025)
by: Orynyak, Igor, et al.
Published: (2025)
Don't Splat your Gaussians: Volumetric Ray-Traced Primitives for Modeling and Rendering Scattering and Emissive Media
by: Condor, Jorge, et al.
Published: (2024)
by: Condor, Jorge, et al.
Published: (2024)
PointRegGPT: Boosting 3D Point Cloud Registration using Generative Point-Cloud Pairs for Training
by: Chen, Suyi, et al.
Published: (2024)
by: Chen, Suyi, et al.
Published: (2024)
Heterogeneous Graph Importance Scoring and Clustering with Automated LLM-based Interpretation
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Consumer Research with Projective Techniques: A Mixed Methods-Focused Review and Empirical Reanalysis
by: France, Stephen L.
Published: (2024)
by: France, Stephen L.
Published: (2024)
TexTile: A Differentiable Metric for Texture Tileability
by: Rodriguez-Pardo, Carlos, et al.
Published: (2024)
by: Rodriguez-Pardo, Carlos, et al.
Published: (2024)
Connected Speech-Based Cognitive Assessment in Chinese and English
by: Luz, Saturnino, et al.
Published: (2024)
by: Luz, Saturnino, et al.
Published: (2024)
Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping
by: Lu, Jingyi, et al.
Published: (2025)
by: Lu, Jingyi, et al.
Published: (2025)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
by: Aristorenas, Aris J.
Published: (2024)
by: Aristorenas, Aris J.
Published: (2024)
Alternative Local Discriminant Bases Using Empirical Expectation and Variance Estimation
by: Fossgaard, Eirik
Published: (1999)
by: Fossgaard, Eirik
Published: (1999)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
by: Kozak, Nazar
Published: (2026)
by: Kozak, Nazar
Published: (2026)
Enhancing Speaker Verification with Whispered Speech via Post-Processing
by: Gołębiowska, Magdalena, et al.
Published: (2026)
by: Gołębiowska, Magdalena, et al.
Published: (2026)
Similar Items
-
Real-World En Call Center Transcripts Dataset with PII Redaction
by: Dao, Ha, et al.
Published: (2025) -
The Concatenator: A Bayesian Approach To Real Time Concatenative Musaicing
by: Tralie, Christopher, et al.
Published: (2024) -
Deep Interest Mining for Intent-Enriched Semantic IDs in Multimodal Generative Recommendation
by: Zeng, Yangchen, et al.
Published: (2026) -
Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling
by: Beskorovainyi, Vladimir
Published: (2026) -
Fine-tuning Pre-trained Audio Models for COVID-19 Detection: A Technical Report
by: de Brito, Daniel Oliveira, et al.
Published: (2025)