Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Anchan, Akshit Pramod, Thomas, Jewelith, Roy, Sritama
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909924631511040
author Anchan, Akshit Pramod
Thomas, Jewelith
Roy, Sritama
author_facet Anchan, Akshit Pramod
Thomas, Jewelith
Roy, Sritama
contents Developing comprehensive assistive technologies requires the seamless integration of visual and auditory perception. This research evaluates the feasibility of a modular architecture inspired by core functionalities of perceptive systems like 'Smart Eye.' We propose and benchmark three independent sensing modules: a Convolutional Neural Network (CNN) for eye state detection (drowsiness/attention), a deep CNN for facial expression recognition, and a Long Short-Term Memory (LSTM) network for voice-based speaker identification. Utilizing the Eyes Image, FER2013, and customized audio datasets, our models achieved accuracies of 93.0%, 97.8%, and 96.89%, respectively. This study demonstrates that lightweight, domain-specific models can achieve high fidelity on discrete tasks, establishing a validated foundation for future real-time, multimodal integration in resource-constrained assistive devices.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20474
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification
Anchan, Akshit Pramod
Thomas, Jewelith
Roy, Sritama
Computer Vision and Pattern Recognition
Machine Learning
68T45
I.2.10; I.2.7; I.5.4
Developing comprehensive assistive technologies requires the seamless integration of visual and auditory perception. This research evaluates the feasibility of a modular architecture inspired by core functionalities of perceptive systems like 'Smart Eye.' We propose and benchmark three independent sensing modules: a Convolutional Neural Network (CNN) for eye state detection (drowsiness/attention), a deep CNN for facial expression recognition, and a Long Short-Term Memory (LSTM) network for voice-based speaker identification. Utilizing the Eyes Image, FER2013, and customized audio datasets, our models achieved accuracies of 93.0%, 97.8%, and 96.89%, respectively. This study demonstrates that lightweight, domain-specific models can achieve high fidelity on discrete tasks, establishing a validated foundation for future real-time, multimodal integration in resource-constrained assistive devices.
title Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification
topic Computer Vision and Pattern Recognition
Machine Learning
68T45
I.2.10; I.2.7; I.5.4
url https://arxiv.org/abs/2511.20474