Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adelson, Trevor, Sethu, Vidhyasaharan, Dang, Ting
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911524395679744
author Adelson, Trevor
Sethu, Vidhyasaharan
Dang, Ting
author_facet Adelson, Trevor
Sethu, Vidhyasaharan
Dang, Ting
contents Deep learning dominates speech processing but relies on massive datasets, global backpropagation-guided weight updates, and produces entangled representations. Assembly Calculus (AC), which models sparse neuronal assemblies via Hebbian plasticity and winner-take-all competition, offers a biologically grounded alternative, yet prior work focused on discrete symbolic inputs. We introduce an AC-based speech processing framework that operates directly on continuous speech by combining three key contributions:(i) neural encoding that converts speech into assembly-compatible spike patterns using probabilistic mel binarisation and population-coded MFCCs; (ii) a multi-area architecture organising assemblies across hierarchical timescales and classes; and (iii) cross-area update schemes for downstream tasks. Applied to two core tasks of boundary detection and segment classification, our framework detects phone (F1=0.69) and word (F1=0.61) boundaries without any weight training, and achieves 47.5% and 45.1% accuracy on phone and command recognition. These results show that AC-based dynamical systems are a viable alternative to deep learning for speech processing.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16923
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
Adelson, Trevor
Sethu, Vidhyasaharan
Dang, Ting
Audio and Speech Processing
Sound
I.2.0; I.2.6; I.2.7
Deep learning dominates speech processing but relies on massive datasets, global backpropagation-guided weight updates, and produces entangled representations. Assembly Calculus (AC), which models sparse neuronal assemblies via Hebbian plasticity and winner-take-all competition, offers a biologically grounded alternative, yet prior work focused on discrete symbolic inputs. We introduce an AC-based speech processing framework that operates directly on continuous speech by combining three key contributions:(i) neural encoding that converts speech into assembly-compatible spike patterns using probabilistic mel binarisation and population-coded MFCCs; (ii) a multi-area architecture organising assemblies across hierarchical timescales and classes; and (iii) cross-area update schemes for downstream tasks. Applied to two core tasks of boundary detection and segment classification, our framework detects phone (F1=0.69) and word (F1=0.61) boundaries without any weight training, and achieves 47.5% and 45.1% accuracy on phone and command recognition. These results show that AC-based dynamical systems are a viable alternative to deep learning for speech processing.
title Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
topic Audio and Speech Processing
Sound
I.2.0; I.2.6; I.2.7
url https://arxiv.org/abs/2603.16923