SyncNet: correlating objective for time delay estimation in audio signals
Fuente:
arXiv
Guardado en:
| Autores principales: | Raina, Akshay, Arora, Vipul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpretable Convolutional SyncNet
por: Park, Sungjoon, et al.
Publicado: (2024)
por: Park, Sungjoon, et al.
Publicado: (2024)
Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions
por: Voran, Stephen D.
Publicado: (2024)
por: Voran, Stephen D.
Publicado: (2024)
Human perception of audio deepfakes: the role of language and speaking style
por: Segundo, Eugenia San, et al.
Publicado: (2025)
por: Segundo, Eugenia San, et al.
Publicado: (2025)
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
por: Kodag, Rahul Bapusaheb, et al.
Publicado: (2025)
por: Kodag, Rahul Bapusaheb, et al.
Publicado: (2025)
Compositional nonlinear audio signal processing with Volterra series
por: Araujo-Simon, Jake
Publicado: (2023)
por: Araujo-Simon, Jake
Publicado: (2023)
Learning from Limited Labels: Transductive Graph Label Propagation for Indian Music Analysis
por: Singh, Parampreet, et al.
Publicado: (2026)
por: Singh, Parampreet, et al.
Publicado: (2026)
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
por: Dutta, Sagar, et al.
Publicado: (2025)
por: Dutta, Sagar, et al.
Publicado: (2025)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
por: Wang, Ziqian, et al.
Publicado: (2025)
por: Wang, Ziqian, et al.
Publicado: (2025)
Synthetic training set generation using text-to-audio models for environmental sound classification
por: Ronchini, Francesca, et al.
Publicado: (2024)
por: Ronchini, Francesca, et al.
Publicado: (2024)
Revisiting proximity effect using broadband signals
por: Millot, Laurent, et al.
Publicado: (2024)
por: Millot, Laurent, et al.
Publicado: (2024)
Benchmarking multi-component signal processing methods in the time-frequency plane
por: Miramont, Juan M., et al.
Publicado: (2024)
por: Miramont, Juan M., et al.
Publicado: (2024)
Unsupervised detection and classification of heartbeats using the dissimilarity matrix in PCG signals
por: Torre-Cruz, J., et al.
Publicado: (2024)
por: Torre-Cruz, J., et al.
Publicado: (2024)
Real time fault detection in 3D printers using Convolutional Neural Networks and acoustic signals
por: Waheed, Muhammad Fasih, et al.
Publicado: (2026)
por: Waheed, Muhammad Fasih, et al.
Publicado: (2026)
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
Sound field estimation with moving microphones using kernel ridge regression
por: Brunnström, Jesper, et al.
Publicado: (2025)
por: Brunnström, Jesper, et al.
Publicado: (2025)
Phase-Only Positioning in Distributed MIMO Under Phase Impairments: AP Selection Using Deep Learning
por: Ayten, Fatih, et al.
Publicado: (2026)
por: Ayten, Fatih, et al.
Publicado: (2026)
FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
por: Choi, Yuseon, et al.
Publicado: (2025)
por: Choi, Yuseon, et al.
Publicado: (2025)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
por: Sato, Hiroshi, et al.
Publicado: (2024)
por: Sato, Hiroshi, et al.
Publicado: (2024)
Audio signal interpolation using optimal transportation of spectrograms
por: Valdivia, David, et al.
Publicado: (2025)
por: Valdivia, David, et al.
Publicado: (2025)
Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
por: Sechaud, Victor, et al.
Publicado: (2024)
por: Sechaud, Victor, et al.
Publicado: (2024)
Using perceptive subbands analysis to perform audio scenes cartography
por: Millot, Laurent, et al.
Publicado: (2024)
por: Millot, Laurent, et al.
Publicado: (2024)
Deep learning classification system for coconut maturity levels based on acoustic signals
por: Caladcad, June Anne, et al.
Publicado: (2024)
por: Caladcad, June Anne, et al.
Publicado: (2024)
Spectrogram features for audio and speech analysis
por: McLoughlin, Ian, et al.
Publicado: (2026)
por: McLoughlin, Ian, et al.
Publicado: (2026)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
por: Yang, Yujie, et al.
Publicado: (2025)
por: Yang, Yujie, et al.
Publicado: (2025)
Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers
por: Silaev, Mikhail, et al.
Publicado: (2026)
por: Silaev, Mikhail, et al.
Publicado: (2026)
Analytical model for the relation between signal bandwidth and spatial resolution in Steered-Response Power Phase Transform (SRP-PHAT) maps
por: Garcia-Barrios, Guillermo, et al.
Publicado: (2024)
por: Garcia-Barrios, Guillermo, et al.
Publicado: (2024)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
por: Yan, Haoyin, et al.
Publicado: (2024)
por: Yan, Haoyin, et al.
Publicado: (2024)
Time-domain sound field estimation using kernel ridge regression
por: Brunnström, Jesper, et al.
Publicado: (2025)
por: Brunnström, Jesper, et al.
Publicado: (2025)
Frequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection
por: Liao, Yuan, et al.
Publicado: (2025)
por: Liao, Yuan, et al.
Publicado: (2025)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
por: Messina, Francisco, et al.
Publicado: (2025)
por: Messina, Francisco, et al.
Publicado: (2025)
Uncertainty Quantification in Melody Estimation using Histogram Representation
por: Saxena, Kavya Ranjan, et al.
Publicado: (2025)
por: Saxena, Kavya Ranjan, et al.
Publicado: (2025)
Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
por: Kodag, Rahul Bapusaheb, et al.
Publicado: (2026)
por: Kodag, Rahul Bapusaheb, et al.
Publicado: (2026)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
por: Hao, Xiang, et al.
Publicado: (2020)
por: Hao, Xiang, et al.
Publicado: (2020)
FlowDec: A flow-based full-band general audio codec with high perceptual quality
por: Welker, Simon, et al.
Publicado: (2025)
por: Welker, Simon, et al.
Publicado: (2025)
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
por: Singh, Anup, et al.
Publicado: (2025)
por: Singh, Anup, et al.
Publicado: (2025)
Improving Active Learning for Melody Estimation by Disentangling Uncertainties
por: Jaiswal, Aayush, et al.
Publicado: (2025)
por: Jaiswal, Aayush, et al.
Publicado: (2025)
SuperM2M: Supervised and Mixture-to-Mixture Co-Learning for Speech Enhancement and Noise-Robust ASR
por: Wang, Zhong-Qiu
Publicado: (2024)
por: Wang, Zhong-Qiu
Publicado: (2024)
Joint Spectrogram Separation and TDOA Estimation using Optimal Transport
por: Fabiani, Linda, et al.
Publicado: (2025)
por: Fabiani, Linda, et al.
Publicado: (2025)
Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process
por: Bologni, Giovanni, et al.
Publicado: (2025)
por: Bologni, Giovanni, et al.
Publicado: (2025)
Parameter-Efficient Fine-Tuning of Foundation Models for CLP Speech Classification
por: Bhattacharjee, Susmita, et al.
Publicado: (2025)
por: Bhattacharjee, Susmita, et al.
Publicado: (2025)
Ejemplares similares
-
Interpretable Convolutional SyncNet
por: Park, Sungjoon, et al.
Publicado: (2024) -
Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions
por: Voran, Stephen D.
Publicado: (2024) -
Human perception of audio deepfakes: the role of language and speaking style
por: Segundo, Eugenia San, et al.
Publicado: (2025) -
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
por: Kodag, Rahul Bapusaheb, et al.
Publicado: (2025) -
Compositional nonlinear audio signal processing with Volterra series
por: Araujo-Simon, Jake
Publicado: (2023)