Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Abavisani, Mahdi, Joze, Hamid Reza Vaezi, Patel, Vishal M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2018
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Deep Multimodal Subspace Clustering Networks
di: Abavisani, Mahdi, et al.
Pubblicazione: (2018)
di: Abavisani, Mahdi, et al.
Pubblicazione: (2018)
Deep Sparse Representation-based Classification
di: Abavisani, Mahdi, et al.
Pubblicazione: (2019)
di: Abavisani, Mahdi, et al.
Pubblicazione: (2019)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
Image Reconstruction as a Tool for Feature Analysis
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)
Interpretable Machine Learning-Derived Spectral Indices for Vegetation Monitoring
di: Lotfi, Ali, et al.
Pubblicazione: (2025)
di: Lotfi, Ali, et al.
Pubblicazione: (2025)
Training a Student Expert via Semi-Supervised Foundation Model Distillation
di: Taghavi, Pardis, et al.
Pubblicazione: (2026)
di: Taghavi, Pardis, et al.
Pubblicazione: (2026)
DeepShade: Enable Shade Simulation by Text-conditioned Image Generation
di: Da, Longchao, et al.
Pubblicazione: (2025)
di: Da, Longchao, et al.
Pubblicazione: (2025)
Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse
di: Meng, Wei
Pubblicazione: (2025)
di: Meng, Wei
Pubblicazione: (2025)
Deep Learning Approaches for Human Action Recognition in Video Data
di: Xie, Yufei
Pubblicazione: (2024)
di: Xie, Yufei
Pubblicazione: (2024)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
A Tractography Analysis Framework Using Diffusion Maps to Study Thalamic Connectivity in Traumatic Brain Injury
di: Sharma, Akul, et al.
Pubblicazione: (2025)
di: Sharma, Akul, et al.
Pubblicazione: (2025)
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
di: Deng, Zhongchen, et al.
Pubblicazione: (2024)
di: Deng, Zhongchen, et al.
Pubblicazione: (2024)
Unpacking Hateful Memes: Presupposed Context and False Claims
di: Cai, Weibin, et al.
Pubblicazione: (2025)
di: Cai, Weibin, et al.
Pubblicazione: (2025)
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases
di: Bell-Navas, Andrés, et al.
Pubblicazione: (2025)
di: Bell-Navas, Andrés, et al.
Pubblicazione: (2025)
Enhancing Human Action Recognition and Violence Detection Through Deep Learning Audiovisual Fusion
di: Janani, Pooya, et al.
Pubblicazione: (2024)
di: Janani, Pooya, et al.
Pubblicazione: (2024)
Learning Association via Track-Detection Matching for Multi-Object Tracking
di: Adžemović, Momir
Pubblicazione: (2025)
di: Adžemović, Momir
Pubblicazione: (2025)
Supervised Learning Has a Necessary Geometric Blind Spot: Theory, Consequences, and Minimal Repair
di: Rajput, Vishal
Pubblicazione: (2026)
di: Rajput, Vishal
Pubblicazione: (2026)
Extraction Of Cumulative Blobs From Dynamic Gestures
di: Naulakha, Rishabh, et al.
Pubblicazione: (2025)
di: Naulakha, Rishabh, et al.
Pubblicazione: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
di: Patel, Urjitkumar, et al.
Pubblicazione: (2025)
di: Patel, Urjitkumar, et al.
Pubblicazione: (2025)
Enhancing Maritime Object Detection in Real-Time with RT-DETR and Data Augmentation
di: Nemati, Nader
Pubblicazione: (2025)
di: Nemati, Nader
Pubblicazione: (2025)
LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool
di: Ma, Yue, et al.
Pubblicazione: (2024)
di: Ma, Yue, et al.
Pubblicazione: (2024)
K-GMRF: Kinetic Gauss-Markov Random Field for First-Principles Covariance Tracking on Lie Groups
di: Li, ZhiMing
Pubblicazione: (2026)
di: Li, ZhiMing
Pubblicazione: (2026)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
di: Tu, Songjun, et al.
Pubblicazione: (2025)
di: Tu, Songjun, et al.
Pubblicazione: (2025)
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
Fast 3D point clouds retrieval for Large-scale 3D Place Recognition
di: Zede, Chahine-Nicolas, et al.
Pubblicazione: (2025)
di: Zede, Chahine-Nicolas, et al.
Pubblicazione: (2025)
Revisiting Energy-Based Model for Out-of-Distribution Detection
di: Wu, Yifan, et al.
Pubblicazione: (2024)
di: Wu, Yifan, et al.
Pubblicazione: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
SemanticFeels: Semantic Labeling during In-Hand Manipulation
di: Khalil, Anas Al Shikh, et al.
Pubblicazione: (2026)
di: Khalil, Anas Al Shikh, et al.
Pubblicazione: (2026)
VDPP: Video Depth Post-Processing for Speed and Scalability
di: Yoon, Daewon, et al.
Pubblicazione: (2026)
di: Yoon, Daewon, et al.
Pubblicazione: (2026)
Parking Space Detection in the City of Granada
di: Luis, Crespo-Orti, et al.
Pubblicazione: (2025)
di: Luis, Crespo-Orti, et al.
Pubblicazione: (2025)
Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
di: Zhang, Lintong, et al.
Pubblicazione: (2025)
di: Zhang, Lintong, et al.
Pubblicazione: (2025)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
di: Liao, Xinyao, et al.
Pubblicazione: (2025)
di: Liao, Xinyao, et al.
Pubblicazione: (2025)
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
di: Montello, Fabio, et al.
Pubblicazione: (2025)
di: Montello, Fabio, et al.
Pubblicazione: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
di: Su, Yuetong, et al.
Pubblicazione: (2025)
di: Su, Yuetong, et al.
Pubblicazione: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
HieraEdgeNet: A Multi-Scale Edge-Enhanced Framework for Automated Pollen Recognition
di: Long, Yuchong, et al.
Pubblicazione: (2025)
di: Long, Yuchong, et al.
Pubblicazione: (2025)
Nonverbal Immediacy Analysis in Education: A Multimodal Computational Model
di: Petković, Uroš, et al.
Pubblicazione: (2024)
di: Petković, Uroš, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Deep Multimodal Subspace Clustering Networks
di: Abavisani, Mahdi, et al.
Pubblicazione: (2018) -
Deep Sparse Representation-based Classification
di: Abavisani, Mahdi, et al.
Pubblicazione: (2019) -
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
di: Zhang, Junbin, et al.
Pubblicazione: (2022) -
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025) -
Image Reconstruction as a Tool for Feature Analysis
di: Allakhverdov, Eduard, et al.
Pubblicazione: (2025)