Unsupervised Out-of-Distribution Dialect Detection with Mahalanobis Distance
Fuente:
arXiv
Salvato in:
| Autori principali: | Das, Sourya Dipta, Vadi, Yash, Unnam, Abhishek, Yadav, Kuldeep |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Measuring Sound Symbolism in Audio-visual Models
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
di: Ortega-Beltrán, Elena, et al.
Pubblicazione: (2024)
di: Ortega-Beltrán, Elena, et al.
Pubblicazione: (2024)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
di: Wu, Yihan, et al.
Pubblicazione: (2024)
di: Wu, Yihan, et al.
Pubblicazione: (2024)
ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
di: Cai, Kevin, et al.
Pubblicazione: (2024)
di: Cai, Kevin, et al.
Pubblicazione: (2024)
Qwen2.5-Omni Technical Report
di: Xu, Jin, et al.
Pubblicazione: (2025)
di: Xu, Jin, et al.
Pubblicazione: (2025)
Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
Voice Pathology Detection Using Phonation
di: Siva, Sri Raksha, et al.
Pubblicazione: (2025)
di: Siva, Sri Raksha, et al.
Pubblicazione: (2025)
From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization
di: Wahida, Farah, et al.
Pubblicazione: (2025)
di: Wahida, Farah, et al.
Pubblicazione: (2025)
Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings
di: Gohari, Mahyar, et al.
Pubblicazione: (2024)
di: Gohari, Mahyar, et al.
Pubblicazione: (2024)
Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
di: Wen, Liuyuan
Pubblicazione: (2024)
di: Wen, Liuyuan
Pubblicazione: (2024)
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
di: Klein, Nicholas, et al.
Pubblicazione: (2025)
di: Klein, Nicholas, et al.
Pubblicazione: (2025)
End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
di: Di Pierno, Andrea, et al.
Pubblicazione: (2025)
di: Di Pierno, Andrea, et al.
Pubblicazione: (2025)
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
di: Jiang, Xilin, et al.
Pubblicazione: (2026)
di: Jiang, Xilin, et al.
Pubblicazione: (2026)
On the Audio Hallucinations in Large Audio-Video Language Models
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
di: Tan, Weiting, et al.
Pubblicazione: (2025)
di: Tan, Weiting, et al.
Pubblicazione: (2025)
Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
di: Hori, Chiori, et al.
Pubblicazione: (2025)
di: Hori, Chiori, et al.
Pubblicazione: (2025)
Taming Data and Transformers for Audio Generation
di: Haji-Ali, Moayed, et al.
Pubblicazione: (2024)
di: Haji-Ali, Moayed, et al.
Pubblicazione: (2024)
Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
di: Jiang, Xilin, et al.
Pubblicazione: (2025)
di: Jiang, Xilin, et al.
Pubblicazione: (2025)
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
di: Min, Anna, et al.
Pubblicazione: (2025)
di: Min, Anna, et al.
Pubblicazione: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
di: Goncalves, Lucas, et al.
Pubblicazione: (2024)
di: Goncalves, Lucas, et al.
Pubblicazione: (2024)
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
di: Guan, Yiwen, et al.
Pubblicazione: (2024)
di: Guan, Yiwen, et al.
Pubblicazione: (2024)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
di: Ye, Zhen, et al.
Pubblicazione: (2026)
di: Ye, Zhen, et al.
Pubblicazione: (2026)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan
di: Saeed, Muhammad Saad, et al.
Pubblicazione: (2024)
di: Saeed, Muhammad Saad, et al.
Pubblicazione: (2024)
Exploring Green AI for Audio Deepfake Detection
di: Saha, Subhajit, et al.
Pubblicazione: (2024)
di: Saha, Subhajit, et al.
Pubblicazione: (2024)
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
di: Bhagtani, Kratika, et al.
Pubblicazione: (2024)
di: Bhagtani, Kratika, et al.
Pubblicazione: (2024)
Dialectal Coverage And Generalization in Arabic Speech Recognition
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2024)
di: Djanibekov, Amirbek, et al.
Pubblicazione: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
di: Adila, Aulia, et al.
Pubblicazione: (2024)
di: Adila, Aulia, et al.
Pubblicazione: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
Robust Active Speaker Detection in Noisy Environments
di: Vasireddy, Siva Sai Nagender, et al.
Pubblicazione: (2024)
di: Vasireddy, Siva Sai Nagender, et al.
Pubblicazione: (2024)
Benchmarking Cross-Domain Audio-Visual Deception Detection
di: Guo, Xiaobao, et al.
Pubblicazione: (2024)
di: Guo, Xiaobao, et al.
Pubblicazione: (2024)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
di: Astrid, Marcella, et al.
Pubblicazione: (2024)
di: Astrid, Marcella, et al.
Pubblicazione: (2024)
SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving
di: Barik, Ayush, et al.
Pubblicazione: (2026)
di: Barik, Ayush, et al.
Pubblicazione: (2026)
Few-shot Acoustic Synthesis with Multimodal Flow Matching
di: Brunetto, Amandine
Pubblicazione: (2026)
di: Brunetto, Amandine
Pubblicazione: (2026)
pycnet-audio: A Python package to support bioacoustics data processing
di: Ruff, Zachary J., et al.
Pubblicazione: (2025)
di: Ruff, Zachary J., et al.
Pubblicazione: (2025)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
di: Choi, Jeongsoo, et al.
Pubblicazione: (2024)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2024)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
di: Ryu, Hyeonggon, et al.
Pubblicazione: (2025)
di: Ryu, Hyeonggon, et al.
Pubblicazione: (2025)
Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
di: Sun, Changchang, et al.
Pubblicazione: (2025)
di: Sun, Changchang, et al.
Pubblicazione: (2025)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
di: Park, Young-Hu, et al.
Pubblicazione: (2025)
di: Park, Young-Hu, et al.
Pubblicazione: (2025)
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
di: Lai, Yung-Hsuan, et al.
Pubblicazione: (2025)
di: Lai, Yung-Hsuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Measuring Sound Symbolism in Audio-visual Models
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024) -
Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
di: Ortega-Beltrán, Elena, et al.
Pubblicazione: (2024) -
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
di: Wu, Yihan, et al.
Pubblicazione: (2024) -
ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
di: Cai, Kevin, et al.
Pubblicazione: (2024) -
Qwen2.5-Omni Technical Report
di: Xu, Jin, et al.
Pubblicazione: (2025)