Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
Fuente:
arXiv
Saved in:
| Main Authors: | Adelson, Trevor, Sethu, Vidhyasaharan, Dang, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2025)
by: Nfissi, Alaa, et al.
Published: (2025)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
by: Salvi, Giampiero
Published: (2024)
by: Salvi, Giampiero
Published: (2024)
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
by: Ferreira, Alexandre R., et al.
Published: (2023)
by: Ferreira, Alexandre R., et al.
Published: (2023)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
by: Khorrami, Khazar, et al.
Published: (2021)
by: Khorrami, Khazar, et al.
Published: (2021)
Iterative Feature Boosting for Explainable Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2024)
by: Nfissi, Alaa, et al.
Published: (2024)
Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2024)
by: Nfissi, Alaa, et al.
Published: (2024)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
by: Dar, Daniyal Kabir, et al.
Published: (2025)
by: Dar, Daniyal Kabir, et al.
Published: (2025)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
by: Cheng, Zhuangfei, et al.
Published: (2025)
by: Cheng, Zhuangfei, et al.
Published: (2025)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025)
by: Birardi, Giuseppe
Published: (2025)
Prevailing Research Areas for Music AI in the Era of Foundation Models
by: Wei, Megan, et al.
Published: (2024)
by: Wei, Megan, et al.
Published: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
by: Mehta, Shivam, et al.
Published: (2025)
by: Mehta, Shivam, et al.
Published: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
by: Mehta, Shivam, et al.
Published: (2025)
by: Mehta, Shivam, et al.
Published: (2025)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
by: Li, Haowen, et al.
Published: (2025)
by: Li, Haowen, et al.
Published: (2025)
NAAQA: A Neural Architecture for Acoustic Question Answering
by: Abdelnour, Jerome, et al.
Published: (2021)
by: Abdelnour, Jerome, et al.
Published: (2021)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
by: Taskooh, Hamidreza Kazemi, et al.
Published: (2026)
by: Taskooh, Hamidreza Kazemi, et al.
Published: (2026)
Quantifying the effect of speech pathology on automatic and human speaker verification
by: Halpern, Bence Mark, et al.
Published: (2024)
by: Halpern, Bence Mark, et al.
Published: (2024)
A Baseline Multimodal Approach to Emotion Recognition in Conversations
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
by: Mehta, Shivam, et al.
Published: (2024)
by: Mehta, Shivam, et al.
Published: (2024)
A Robust Classification Method using Hybrid Word Embedding for Early Diagnosis of Alzheimer's Disease
by: Li, Yangyang
Published: (2025)
by: Li, Yangyang
Published: (2025)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
by: Bhadra, Dipayan, et al.
Published: (2025)
by: Bhadra, Dipayan, et al.
Published: (2025)
Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
by: Kuhn, Korbinian, et al.
Published: (2024)
by: Kuhn, Korbinian, et al.
Published: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
by: Ahn, Taekyung, et al.
Published: (2024)
by: Ahn, Taekyung, et al.
Published: (2024)
Everyday Speech in the Indian Subcontinent
by: P, Utkarsh
Published: (2024)
by: P, Utkarsh
Published: (2024)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Measuring the Accuracy of Automatic Speech Recognition Solutions
by: Kuhn, Korbinian, et al.
Published: (2024)
by: Kuhn, Korbinian, et al.
Published: (2024)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
by: Robertson, Sean, et al.
Published: (2023)
by: Robertson, Sean, et al.
Published: (2023)
An Automatic Text Classification Method Based on Hierarchical Taxonomies, Neural Networks and Document Embedding: The NETHIC Tool
by: Lomasto, Luigi, et al.
Published: (2026)
by: Lomasto, Luigi, et al.
Published: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
by: Wang, Hsuan-Yu, et al.
Published: (2025)
by: Wang, Hsuan-Yu, et al.
Published: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
by: Viswanathan, Janaki, et al.
Published: (2025)
by: Viswanathan, Janaki, et al.
Published: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
by: Alamr, Meshal, et al.
Published: (2026)
by: Alamr, Meshal, et al.
Published: (2026)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
by: Jain, Sarthak, et al.
Published: (2024)
by: Jain, Sarthak, et al.
Published: (2024)
Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
by: Räsänen, Okko
Published: (2026)
by: Räsänen, Okko
Published: (2026)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
by: Sharma, Manali, et al.
Published: (2026)
by: Sharma, Manali, et al.
Published: (2026)
Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates
by: Kaplanski, Pawel
Published: (2026)
by: Kaplanski, Pawel
Published: (2026)
Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
by: Ackermann, Richard, et al.
Published: (2025)
by: Ackermann, Richard, et al.
Published: (2025)
Similar Items
-
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2025) -
Developing Acoustic Models for Automatic Speech Recognition in Swedish
by: Salvi, Giampiero
Published: (2024) -
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
by: Ferreira, Alexandre R., et al.
Published: (2023) -
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
by: Khorrami, Khazar, et al.
Published: (2021) -
Iterative Feature Boosting for Explainable Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2024)