Domain Adaptation of the Pyannote Diarization Pipeline for Conversational Indonesian Audio
Fuente:
arXiv
Saved in:
| Main Authors: | Prasetyo, Muhammad Daffa'i Rafi, Putra, Ramadhan Andika, Ilmi, Zaidan Naufal, Azizah, Kurniawati |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
by: Chowdhury, Aurchi, et al.
Published: (2026)
by: Chowdhury, Aurchi, et al.
Published: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
by: Wang, Hsuan-Yu, et al.
Published: (2025)
by: Wang, Hsuan-Yu, et al.
Published: (2025)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
by: Adila, Aulia, et al.
Published: (2024)
by: Adila, Aulia, et al.
Published: (2024)
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset
by: Marie, Ambre, et al.
Published: (2025)
by: Marie, Ambre, et al.
Published: (2025)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
by: Chowdhury, MD. Sagor, et al.
Published: (2026)
by: Chowdhury, MD. Sagor, et al.
Published: (2026)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
by: Bhadra, Dipayan, et al.
Published: (2025)
by: Bhadra, Dipayan, et al.
Published: (2025)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
by: Feng, Ruitao, et al.
Published: (2025)
by: Feng, Ruitao, et al.
Published: (2025)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
by: Zaragozá, Lucía Gómez, et al.
Published: (2024)
by: Zaragozá, Lucía Gómez, et al.
Published: (2024)
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
by: Xu, Tianyu, et al.
Published: (2026)
by: Xu, Tianyu, et al.
Published: (2026)
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
by: Christop, Iwona, et al.
Published: (2026)
by: Christop, Iwona, et al.
Published: (2026)
Multi-level SSL Feature Gating for Audio Deepfake Detection
by: Tran, Hoan My, et al.
Published: (2025)
by: Tran, Hoan My, et al.
Published: (2025)
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
by: Cao, Songjun, et al.
Published: (2026)
by: Cao, Songjun, et al.
Published: (2026)
Dual-Model Prediction of Affective Engagement and Vocal Attractiveness from Speaker Expressiveness in Video Learning
by: Suen, Hung-Yue, et al.
Published: (2026)
by: Suen, Hung-Yue, et al.
Published: (2026)
Dereverberation Using Binary Residual Masking with Time-Domain Consistency
by: Williams, Daniel G.
Published: (2025)
by: Williams, Daniel G.
Published: (2025)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
by: Cheng, Zhuangfei, et al.
Published: (2025)
by: Cheng, Zhuangfei, et al.
Published: (2025)
PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
by: Halpern, Bence Mark, et al.
Published: (2026)
by: Halpern, Bence Mark, et al.
Published: (2026)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
by: Dua, Karan, et al.
Published: (2025)
by: Dua, Karan, et al.
Published: (2025)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
by: Sharma, Manali, et al.
Published: (2026)
by: Sharma, Manali, et al.
Published: (2026)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
by: Simões, Lucca Emmanuel Pineli, et al.
Published: (2024)
by: Simões, Lucca Emmanuel Pineli, et al.
Published: (2024)
Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Quantization for OpenAI's Whisper Models: A Comparative Analysis
by: Andreyev, Allison
Published: (2025)
by: Andreyev, Allison
Published: (2025)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Domain-Incremental Continual Learning for Robust and Efficient Keyword Spotting in Resource Constrained Systems
by: Dhungana, Prakash, et al.
Published: (2026)
by: Dhungana, Prakash, et al.
Published: (2026)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
by: Gausen, Anna, et al.
Published: (2026)
by: Gausen, Anna, et al.
Published: (2026)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
by: Wang, Shaowen, et al.
Published: (2025)
by: Wang, Shaowen, et al.
Published: (2025)
A note on short minimal codes from subgeometries
by: Adriaensen, Sam, et al.
Published: (2025)
by: Adriaensen, Sam, et al.
Published: (2025)
Random trade timing and power-law tails in realized prices
by: Seo, Won-Ki
Published: (2026)
by: Seo, Won-Ki
Published: (2026)
Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
by: Xu, Tianyu, et al.
Published: (2025)
by: Xu, Tianyu, et al.
Published: (2025)
Domain Adaptation for Contrastive Audio-Language Models
by: Deshmukh, Soham, et al.
Published: (2024)
by: Deshmukh, Soham, et al.
Published: (2024)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
by: Li, Pengcheng, et al.
Published: (2024)
by: Li, Pengcheng, et al.
Published: (2024)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
by: Woodard, Brandon, et al.
Published: (2025)
by: Woodard, Brandon, et al.
Published: (2025)
Predicting When to Trust Vision-Language Models for Spatial Reasoning
by: Imran, Muhammad, et al.
Published: (2026)
by: Imran, Muhammad, et al.
Published: (2026)
Real-Time Band-Grouped Vocal Denoising Using Sigmoid-Driven Ideal Ratio Masking
by: Williams, Daniel
Published: (2026)
by: Williams, Daniel
Published: (2026)
Stuttering-Aware Automatic Speech Recognition for Indonesian Language
by: Muhammad, Fadhil, et al.
Published: (2026)
by: Muhammad, Fadhil, et al.
Published: (2026)
Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
by: Zhang, Lin, et al.
Published: (2024)
by: Zhang, Lin, et al.
Published: (2024)
BlasBench: An Open Benchmark for Irish Speech Recognition
by: Raj, Jyoutir, et al.
Published: (2026)
by: Raj, Jyoutir, et al.
Published: (2026)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
by: He, Mao-Kui, et al.
Published: (2024)
by: He, Mao-Kui, et al.
Published: (2024)
Taming Audio VAEs via Target-KL Regularization
by: Seetharaman, Prem, et al.
Published: (2026)
by: Seetharaman, Prem, et al.
Published: (2026)
Similar Items
-
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
by: Chowdhury, Aurchi, et al.
Published: (2026) -
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
by: Wang, Hsuan-Yu, et al.
Published: (2025) -
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
by: Adila, Aulia, et al.
Published: (2024) -
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset
by: Marie, Ambre, et al.
Published: (2025) -
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
by: Chowdhury, MD. Sagor, et al.
Published: (2026)