Gespeichert in:
| Hauptverfasser: | Zeng, Chunyan, Zhao, Yuhao, Wang, Zhifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2411.03668 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
von: Ge, Zirui, et al.
Veröffentlicht: (2023)
von: Ge, Zirui, et al.
Veröffentlicht: (2023)
Self-Supervised Learning of Spatial Acoustic Representation with Cross-Channel Signal Reconstruction and Multi-Channel Conformer
von: Yang, Bing, et al.
Veröffentlicht: (2023)
von: Yang, Bing, et al.
Veröffentlicht: (2023)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings
von: Wu, Fan, et al.
Veröffentlicht: (2026)
von: Wu, Fan, et al.
Veröffentlicht: (2026)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings
von: Zhao, He, et al.
Veröffentlicht: (2024)
von: Zhao, He, et al.
Veröffentlicht: (2024)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
von: Ding, Jiani, et al.
Veröffentlicht: (2025)
von: Ding, Jiani, et al.
Veröffentlicht: (2025)
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
von: Halim, Jule Valendo, et al.
Veröffentlicht: (2025)
von: Halim, Jule Valendo, et al.
Veröffentlicht: (2025)
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
Learnable Pulse Accumulation for On-Device Speech Recognition: How Much Attention Do You Need?
von: Shkolnikov, Yakov Pyotr
Veröffentlicht: (2026)
von: Shkolnikov, Yakov Pyotr
Veröffentlicht: (2026)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2024)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
BickGraphing: Web-Based Application for Visual Inspection of Audio Recordings
von: Seow, Kayley, et al.
Veröffentlicht: (2026)
von: Seow, Kayley, et al.
Veröffentlicht: (2026)
Lightweight Resolution-Aware Audio Deepfake Detection via Cross-Scale Attention and Consistency Learning
von: Shahriar, K. A.
Veröffentlicht: (2026)
von: Shahriar, K. A.
Veröffentlicht: (2026)
Multi-Channel Acoustic Echo Cancellation Based on Direction-of-Arrival Estimation
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Few-Shot Bioacoustic Event Detection with Frame-Level Embedding Learning System
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
Water Flow Detection Device Based on Sound Data Analysis and Machine Learning to Detect Water Leakage
von: Pourmehrani, Hossein, et al.
Veröffentlicht: (2025)
von: Pourmehrani, Hossein, et al.
Veröffentlicht: (2025)
Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
von: Tan, Chao, et al.
Veröffentlicht: (2024)
von: Tan, Chao, et al.
Veröffentlicht: (2024)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024) -
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024) -
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
von: Gao, Yingxue, et al.
Veröffentlicht: (2024) -
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026) -
Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
von: Ge, Zirui, et al.
Veröffentlicht: (2023)