Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Anchan, Akshit Pramod, Thomas, Jewelith, Roy, Sritama |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
by: Adžemović, Momir
Published: (2025)
by: Adžemović, Momir
Published: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
by: Sáez, Arnau Igualde, et al.
Published: (2025)
by: Sáez, Arnau Igualde, et al.
Published: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
by: Tu, Songjun, et al.
Published: (2025)
by: Tu, Songjun, et al.
Published: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025)
by: Zinnen, Mathias, et al.
Published: (2025)
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
by: Michalkiewicz, Mateusz, et al.
Published: (2025)
by: Michalkiewicz, Mateusz, et al.
Published: (2025)
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
by: Hua, Wenjie, et al.
Published: (2025)
by: Hua, Wenjie, et al.
Published: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
Semantically Guided Adversarial Testing of Vision Models Using Language Models
by: Filus, Katarzyna, et al.
Published: (2025)
by: Filus, Katarzyna, et al.
Published: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
Predicting When to Trust Vision-Language Models for Spatial Reasoning
by: Imran, Muhammad, et al.
Published: (2026)
by: Imran, Muhammad, et al.
Published: (2026)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
by: Borisov, Vadim
Published: (2026)
by: Borisov, Vadim
Published: (2026)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Geometric-Stochastic Multimodal Deep Learning for Predictive Modeling of SUDEP and Stroke Vulnerability
by: Girish, Preksha, et al.
Published: (2025)
by: Girish, Preksha, et al.
Published: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
by: Patel, Urjitkumar, et al.
Published: (2025)
by: Patel, Urjitkumar, et al.
Published: (2025)
Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
by: Pandey, Anupam, et al.
Published: (2025)
by: Pandey, Anupam, et al.
Published: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
by: Farzulla, Murad
Published: (2026)
by: Farzulla, Murad
Published: (2026)
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
by: Cao, Songjun, et al.
Published: (2026)
by: Cao, Songjun, et al.
Published: (2026)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
by: Huo, Dongjie, et al.
Published: (2026)
by: Huo, Dongjie, et al.
Published: (2026)
Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Prediction
by: Artru, Noé, et al.
Published: (2026)
by: Artru, Noé, et al.
Published: (2026)
MVTamperBench: Evaluating Robustness of Vision-Language Models
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
The MSR-Video to Text Dataset with Clean Annotations
by: Chen, Haoran, et al.
Published: (2021)
by: Chen, Haoran, et al.
Published: (2021)
HieraEdgeNet: A Multi-Scale Edge-Enhanced Framework for Automated Pollen Recognition
by: Long, Yuchong, et al.
Published: (2025)
by: Long, Yuchong, et al.
Published: (2025)
TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
by: Nauen, Tobias Christian, et al.
Published: (2024)
by: Nauen, Tobias Christian, et al.
Published: (2024)
Motion-Based Sign Language Video Summarization using Curvature and Torsion
by: Sartinas, Evangelos G., et al.
Published: (2023)
by: Sartinas, Evangelos G., et al.
Published: (2023)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
by: Yang, Baoyao, et al.
Published: (2025)
by: Yang, Baoyao, et al.
Published: (2025)
Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
by: Nasriddinov, Firdavs, et al.
Published: (2025)
by: Nasriddinov, Firdavs, et al.
Published: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
by: Yanambakkam, Hemanth Teja, et al.
Published: (2025)
by: Yanambakkam, Hemanth Teja, et al.
Published: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
by: Lyu, Wenbo, et al.
Published: (2025)
by: Lyu, Wenbo, et al.
Published: (2025)
Salient Concept-Aware Generative Data Augmentation
by: Zhao, Tianchen, et al.
Published: (2025)
by: Zhao, Tianchen, et al.
Published: (2025)
Fast 3D point clouds retrieval for Large-scale 3D Place Recognition
by: Zede, Chahine-Nicolas, et al.
Published: (2025)
by: Zede, Chahine-Nicolas, et al.
Published: (2025)
An M-Health Algorithmic Approach to Identify and Assess Physiotherapy Exercises in Real Time
by: Kandylakis, Stylianos, et al.
Published: (2025)
by: Kandylakis, Stylianos, et al.
Published: (2025)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
by: Dong, Zixuan, et al.
Published: (2025)
by: Dong, Zixuan, et al.
Published: (2025)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
by: Qi, Dekang, et al.
Published: (2026)
by: Qi, Dekang, et al.
Published: (2026)
Non-Robust Features are Not Always Useful in One-Class Classification
by: Lau, Matthew, et al.
Published: (2024)
by: Lau, Matthew, et al.
Published: (2024)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
by: Bell-Navas, Andrés, et al.
Published: (2025)
by: Bell-Navas, Andrés, et al.
Published: (2025)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
by: Asanuma, Haruka, et al.
Published: (2025)
by: Asanuma, Haruka, et al.
Published: (2025)
Similar Items
-
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025) -
Learning Association via Track-Detection Matching for Multi-Object Tracking
by: Adžemović, Momir
Published: (2025) -
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
by: Sáez, Arnau Igualde, et al.
Published: (2025) -
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
by: Tu, Songjun, et al.
Published: (2025) -
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025)