SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nfissi, Alaa, Bouachir, Wassim, Bouguila, Nizar, Mishara, Brian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Iterative Feature Boosting for Explainable Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
A Baseline Multimodal Approach to Emotion Recognition in Conversations
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
von: Salvi, Giampiero
Veröffentlicht: (2024)
von: Salvi, Giampiero
Veröffentlicht: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
Measuring the Accuracy of Automatic Speech Recognition Solutions
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
Quantifying the effect of speech pathology on automatic and human speaker verification
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)
von: Li, Haowen, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
A Voice-based Triage for Type 2 Diabetes using a Conversational Virtual Assistant in the Home Environment
von: Summoogum, Kelvin, et al.
Veröffentlicht: (2024)
von: Summoogum, Kelvin, et al.
Veröffentlicht: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Everyday Speech in the Indian Subcontinent
von: P, Utkarsh
Veröffentlicht: (2024)
von: P, Utkarsh
Veröffentlicht: (2024)
Prevailing Research Areas for Music AI in the Era of Foundation Models
von: Wei, Megan, et al.
Veröffentlicht: (2024)
von: Wei, Megan, et al.
Veröffentlicht: (2024)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
von: Ferreira, Alexandre R., et al.
Veröffentlicht: (2023)
von: Ferreira, Alexandre R., et al.
Veröffentlicht: (2023)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
NAAQA: A Neural Architecture for Acoustic Question Answering
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
von: Kozak, Nazar
Veröffentlicht: (2026)
von: Kozak, Nazar
Veröffentlicht: (2026)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Passive Underwater Acoustic Signal Separation based on Feature Decoupling Dual-path Network
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
von: Gautam, Sushant, et al.
Veröffentlicht: (2024)
von: Gautam, Sushant, et al.
Veröffentlicht: (2024)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Iterative Feature Boosting for Explainable Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024) -
Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024) -
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026) -
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025) -
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)