Sanvaad: A Multimodal Accessibility Framework for ISL Recognition and Voice-Based Interaction
Fuente:
arXiv
Guardado en:
| Autores principales: | Revankar, Kush, Deshpande, Shreyas, Sayeed, Araham, Tandale, Ansh, Bobde, Sarika |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scale-Aware Recognition in Satellite Images under Resource Constraints
por: Revankar, Shreelekha, et al.
Publicado: (2024)
por: Revankar, Shreelekha, et al.
Publicado: (2024)
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
por: Revankar, Shreelekha, et al.
Publicado: (2025)
por: Revankar, Shreelekha, et al.
Publicado: (2025)
Neural Radiance Fields: Past, Present, and Future
por: Mittal, Ansh
Publicado: (2023)
por: Mittal, Ansh
Publicado: (2023)
A Comparison of Lightweight Deep Learning Models for Particulate-Matter Nowcasting in the Indian Subcontinent & Surrounding Regions
por: Kushwaha, Ansh, et al.
Publicado: (2025)
por: Kushwaha, Ansh, et al.
Publicado: (2025)
Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
por: Bohy, Hugo, et al.
Publicado: (2025)
por: Bohy, Hugo, et al.
Publicado: (2025)
ViKANformer: Embedding Kolmogorov Arnold Networks in Vision Transformers for Pattern-Based Learning
por: S, Shreyas, et al.
Publicado: (2025)
por: S, Shreyas, et al.
Publicado: (2025)
Interpretable Underwater Diver Gesture Recognition
por: Mangalvedhekar, Sudeep, et al.
Publicado: (2023)
por: Mangalvedhekar, Sudeep, et al.
Publicado: (2023)
IFSENet : Harnessing Sparse Iterations for Interactive Few-shot Segmentation Excellence
por: Chandgothia, Shreyas, et al.
Publicado: (2024)
por: Chandgothia, Shreyas, et al.
Publicado: (2024)
SHARE: Single-view Human Adversarial REconstruction
por: Revankar, Shreelekha, et al.
Publicado: (2023)
por: Revankar, Shreelekha, et al.
Publicado: (2023)
Rethinking the Threat and Accessibility of Adversarial Attacks against Face Recognition Systems
por: Cao, Yuxin, et al.
Publicado: (2024)
por: Cao, Yuxin, et al.
Publicado: (2024)
Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
por: Liu, Huimin, et al.
Publicado: (2025)
por: Liu, Huimin, et al.
Publicado: (2025)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
por: Zhang, Haoyang, et al.
Publicado: (2025)
por: Zhang, Haoyang, et al.
Publicado: (2025)
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization
por: Anshul, Ashutosh, et al.
Publicado: (2025)
por: Anshul, Ashutosh, et al.
Publicado: (2025)
UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
por: Qi, Zhangshuo, et al.
Publicado: (2025)
por: Qi, Zhangshuo, et al.
Publicado: (2025)
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
por: Ziaeetabar, Fatemeh, et al.
Publicado: (2025)
por: Ziaeetabar, Fatemeh, et al.
Publicado: (2025)
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
por: Romero, Julia Lee, et al.
Publicado: (2025)
por: Romero, Julia Lee, et al.
Publicado: (2025)
SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
por: Liu, Hongda, et al.
Publicado: (2025)
por: Liu, Hongda, et al.
Publicado: (2025)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
por: Wang, Mengmeng, et al.
Publicado: (2024)
por: Wang, Mengmeng, et al.
Publicado: (2024)
SuperEx: Enhancing Indoor Mapping and Exploration using Non-Line-of-Sight Perception
por: Garg, Kush, et al.
Publicado: (2025)
por: Garg, Kush, et al.
Publicado: (2025)
Connecting the Dots: Leveraging Spatio-Temporal Graph Neural Networks for Accurate Bangla Sign Language Recognition
por: Shahgir, Haz Sameen, et al.
Publicado: (2024)
por: Shahgir, Haz Sameen, et al.
Publicado: (2024)
Explicit Interaction for Fusion-Based Place Recognition
por: Xu, Jingyi, et al.
Publicado: (2024)
por: Xu, Jingyi, et al.
Publicado: (2024)
A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product
por: Xiang, Ao, et al.
Publicado: (2024)
por: Xiang, Ao, et al.
Publicado: (2024)
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
por: Deng, Xiang, et al.
Publicado: (2026)
por: Deng, Xiang, et al.
Publicado: (2026)
A Trustworthy Method for Multimodal Emotion Recognition
por: Xue, Junxiao, et al.
Publicado: (2025)
por: Xue, Junxiao, et al.
Publicado: (2025)
Real-Time Detection and Analysis of Vehicles and Pedestrians using Deep Learning
por: Sadik, Md Nahid, et al.
Publicado: (2024)
por: Sadik, Md Nahid, et al.
Publicado: (2024)
Towards Visual Syntactical Understanding
por: Chowdhury, Sayeed Shafayet, et al.
Publicado: (2024)
por: Chowdhury, Sayeed Shafayet, et al.
Publicado: (2024)
Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks
por: Karri, Sai Likhith, et al.
Publicado: (2025)
por: Karri, Sai Likhith, et al.
Publicado: (2025)
TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition
por: Yin, Wen, et al.
Publicado: (2025)
por: Yin, Wen, et al.
Publicado: (2025)
Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model
por: Ge, Mengying, et al.
Publicado: (2024)
por: Ge, Mengying, et al.
Publicado: (2024)
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
por: Kim, Junhyeok, et al.
Publicado: (2025)
por: Kim, Junhyeok, et al.
Publicado: (2025)
in-Car Biometrics (iCarB) Datasets for Driver Recognition: Face, Fingerprint, and Voice
por: Hahn, Vedrana Krivokuca, et al.
Publicado: (2024)
por: Hahn, Vedrana Krivokuca, et al.
Publicado: (2024)
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
por: Phung, Minh-Chi, et al.
Publicado: (2026)
por: Phung, Minh-Chi, et al.
Publicado: (2026)
Adaptive Sensitivity Analysis for Robust Augmentation against Natural Corruptions in Image Segmentation
por: Zheng, Laura, et al.
Publicado: (2024)
por: Zheng, Laura, et al.
Publicado: (2024)
Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition
por: Liu, Ran, et al.
Publicado: (2025)
por: Liu, Ran, et al.
Publicado: (2025)
Voice-Assisted Real-Time Traffic Sign Recognition System Using Convolutional Neural Network
por: Manawadu, Mayura, et al.
Publicado: (2024)
por: Manawadu, Mayura, et al.
Publicado: (2024)
ZeBROD: Zero-Retraining Based Recognition and Object Detection Framework
por: Hidayatullah, Priyanto, et al.
Publicado: (2025)
por: Hidayatullah, Priyanto, et al.
Publicado: (2025)
Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder
por: Peng, Chong, et al.
Publicado: (2024)
por: Peng, Chong, et al.
Publicado: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
por: Song, Yehun, et al.
Publicado: (2025)
por: Song, Yehun, et al.
Publicado: (2025)
Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline
por: Wei, Tianqi, et al.
Publicado: (2024)
por: Wei, Tianqi, et al.
Publicado: (2024)
Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
por: Li, Yong, et al.
Publicado: (2026)
por: Li, Yong, et al.
Publicado: (2026)
Ejemplares similares
-
Scale-Aware Recognition in Satellite Images under Resource Constraints
por: Revankar, Shreelekha, et al.
Publicado: (2024) -
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
por: Revankar, Shreelekha, et al.
Publicado: (2025) -
Neural Radiance Fields: Past, Present, and Future
por: Mittal, Ansh
Publicado: (2023) -
A Comparison of Lightweight Deep Learning Models for Particulate-Matter Nowcasting in the Indian Subcontinent & Surrounding Regions
por: Kushwaha, Ansh, et al.
Publicado: (2025) -
Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
por: Bohy, Hugo, et al.
Publicado: (2025)