Sanvaad: A Multimodal Accessibility Framework for ISL Recognition and Voice-Based Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Revankar, Kush, Deshpande, Shreyas, Sayeed, Araham, Tandale, Ansh, Bobde, Sarika |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scale-Aware Recognition in Satellite Images under Resource Constraints
by: Revankar, Shreelekha, et al.
Published: (2024)
by: Revankar, Shreelekha, et al.
Published: (2024)
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
by: Revankar, Shreelekha, et al.
Published: (2025)
by: Revankar, Shreelekha, et al.
Published: (2025)
Neural Radiance Fields: Past, Present, and Future
by: Mittal, Ansh
Published: (2023)
by: Mittal, Ansh
Published: (2023)
A Comparison of Lightweight Deep Learning Models for Particulate-Matter Nowcasting in the Indian Subcontinent & Surrounding Regions
by: Kushwaha, Ansh, et al.
Published: (2025)
by: Kushwaha, Ansh, et al.
Published: (2025)
Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
by: Bohy, Hugo, et al.
Published: (2025)
by: Bohy, Hugo, et al.
Published: (2025)
ViKANformer: Embedding Kolmogorov Arnold Networks in Vision Transformers for Pattern-Based Learning
by: S, Shreyas, et al.
Published: (2025)
by: S, Shreyas, et al.
Published: (2025)
Interpretable Underwater Diver Gesture Recognition
by: Mangalvedhekar, Sudeep, et al.
Published: (2023)
by: Mangalvedhekar, Sudeep, et al.
Published: (2023)
IFSENet : Harnessing Sparse Iterations for Interactive Few-shot Segmentation Excellence
by: Chandgothia, Shreyas, et al.
Published: (2024)
by: Chandgothia, Shreyas, et al.
Published: (2024)
SHARE: Single-view Human Adversarial REconstruction
by: Revankar, Shreelekha, et al.
Published: (2023)
by: Revankar, Shreelekha, et al.
Published: (2023)
Rethinking the Threat and Accessibility of Adversarial Attacks against Face Recognition Systems
by: Cao, Yuxin, et al.
Published: (2024)
by: Cao, Yuxin, et al.
Published: (2024)
Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
by: Liu, Huimin, et al.
Published: (2025)
by: Liu, Huimin, et al.
Published: (2025)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
by: Zhang, Haoyang, et al.
Published: (2025)
by: Zhang, Haoyang, et al.
Published: (2025)
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization
by: Anshul, Ashutosh, et al.
Published: (2025)
by: Anshul, Ashutosh, et al.
Published: (2025)
UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
by: Qi, Zhangshuo, et al.
Published: (2025)
by: Qi, Zhangshuo, et al.
Published: (2025)
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
by: Ziaeetabar, Fatemeh, et al.
Published: (2025)
by: Ziaeetabar, Fatemeh, et al.
Published: (2025)
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
by: Romero, Julia Lee, et al.
Published: (2025)
by: Romero, Julia Lee, et al.
Published: (2025)
SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
by: Liu, Hongda, et al.
Published: (2025)
by: Liu, Hongda, et al.
Published: (2025)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
by: Wang, Mengmeng, et al.
Published: (2024)
by: Wang, Mengmeng, et al.
Published: (2024)
SuperEx: Enhancing Indoor Mapping and Exploration using Non-Line-of-Sight Perception
by: Garg, Kush, et al.
Published: (2025)
by: Garg, Kush, et al.
Published: (2025)
Connecting the Dots: Leveraging Spatio-Temporal Graph Neural Networks for Accurate Bangla Sign Language Recognition
by: Shahgir, Haz Sameen, et al.
Published: (2024)
by: Shahgir, Haz Sameen, et al.
Published: (2024)
Explicit Interaction for Fusion-Based Place Recognition
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product
by: Xiang, Ao, et al.
Published: (2024)
by: Xiang, Ao, et al.
Published: (2024)
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
by: Deng, Xiang, et al.
Published: (2026)
by: Deng, Xiang, et al.
Published: (2026)
A Trustworthy Method for Multimodal Emotion Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
Real-Time Detection and Analysis of Vehicles and Pedestrians using Deep Learning
by: Sadik, Md Nahid, et al.
Published: (2024)
by: Sadik, Md Nahid, et al.
Published: (2024)
Towards Visual Syntactical Understanding
by: Chowdhury, Sayeed Shafayet, et al.
Published: (2024)
by: Chowdhury, Sayeed Shafayet, et al.
Published: (2024)
Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks
by: Karri, Sai Likhith, et al.
Published: (2025)
by: Karri, Sai Likhith, et al.
Published: (2025)
TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition
by: Yin, Wen, et al.
Published: (2025)
by: Yin, Wen, et al.
Published: (2025)
Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model
by: Ge, Mengying, et al.
Published: (2024)
by: Ge, Mengying, et al.
Published: (2024)
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
in-Car Biometrics (iCarB) Datasets for Driver Recognition: Face, Fingerprint, and Voice
by: Hahn, Vedrana Krivokuca, et al.
Published: (2024)
by: Hahn, Vedrana Krivokuca, et al.
Published: (2024)
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
by: Phung, Minh-Chi, et al.
Published: (2026)
by: Phung, Minh-Chi, et al.
Published: (2026)
Adaptive Sensitivity Analysis for Robust Augmentation against Natural Corruptions in Image Segmentation
by: Zheng, Laura, et al.
Published: (2024)
by: Zheng, Laura, et al.
Published: (2024)
Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition
by: Liu, Ran, et al.
Published: (2025)
by: Liu, Ran, et al.
Published: (2025)
Voice-Assisted Real-Time Traffic Sign Recognition System Using Convolutional Neural Network
by: Manawadu, Mayura, et al.
Published: (2024)
by: Manawadu, Mayura, et al.
Published: (2024)
ZeBROD: Zero-Retraining Based Recognition and Object Detection Framework
by: Hidayatullah, Priyanto, et al.
Published: (2025)
by: Hidayatullah, Priyanto, et al.
Published: (2025)
Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder
by: Peng, Chong, et al.
Published: (2024)
by: Peng, Chong, et al.
Published: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
by: Song, Yehun, et al.
Published: (2025)
by: Song, Yehun, et al.
Published: (2025)
Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline
by: Wei, Tianqi, et al.
Published: (2024)
by: Wei, Tianqi, et al.
Published: (2024)
Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
by: Li, Yong, et al.
Published: (2026)
by: Li, Yong, et al.
Published: (2026)
Similar Items
-
Scale-Aware Recognition in Satellite Images under Resource Constraints
by: Revankar, Shreelekha, et al.
Published: (2024) -
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
by: Revankar, Shreelekha, et al.
Published: (2025) -
Neural Radiance Fields: Past, Present, and Future
by: Mittal, Ansh
Published: (2023) -
A Comparison of Lightweight Deep Learning Models for Particulate-Matter Nowcasting in the Indian Subcontinent & Surrounding Regions
by: Kushwaha, Ansh, et al.
Published: (2025) -
Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
by: Bohy, Hugo, et al.
Published: (2025)