Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909924631511040 |
|---|---|
| author | Anchan, Akshit Pramod Thomas, Jewelith Roy, Sritama |
| author_facet | Anchan, Akshit Pramod Thomas, Jewelith Roy, Sritama |
| contents | Developing comprehensive assistive technologies requires the seamless integration of visual and auditory perception. This research evaluates the feasibility of a modular architecture inspired by core functionalities of perceptive systems like 'Smart Eye.' We propose and benchmark three independent sensing modules: a Convolutional Neural Network (CNN) for eye state detection (drowsiness/attention), a deep CNN for facial expression recognition, and a Long Short-Term Memory (LSTM) network for voice-based speaker identification. Utilizing the Eyes Image, FER2013, and customized audio datasets, our models achieved accuracies of 93.0%, 97.8%, and 96.89%, respectively. This study demonstrates that lightweight, domain-specific models can achieve high fidelity on discrete tasks, establishing a validated foundation for future real-time, multimodal integration in resource-constrained assistive devices. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_20474 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification Anchan, Akshit Pramod Thomas, Jewelith Roy, Sritama Computer Vision and Pattern Recognition Machine Learning 68T45 I.2.10; I.2.7; I.5.4 Developing comprehensive assistive technologies requires the seamless integration of visual and auditory perception. This research evaluates the feasibility of a modular architecture inspired by core functionalities of perceptive systems like 'Smart Eye.' We propose and benchmark three independent sensing modules: a Convolutional Neural Network (CNN) for eye state detection (drowsiness/attention), a deep CNN for facial expression recognition, and a Long Short-Term Memory (LSTM) network for voice-based speaker identification. Utilizing the Eyes Image, FER2013, and customized audio datasets, our models achieved accuracies of 93.0%, 97.8%, and 96.89%, respectively. This study demonstrates that lightweight, domain-specific models can achieve high fidelity on discrete tasks, establishing a validated foundation for future real-time, multimodal integration in resource-constrained assistive devices. |
| title | Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification |
| topic | Computer Vision and Pattern Recognition Machine Learning 68T45 I.2.10; I.2.7; I.5.4 |
| url | https://arxiv.org/abs/2511.20474 |