Continuous Sign Language Recognition System using Deep Learning with MediaPipe Holistic
Fuente:
arXiv
Saved in:
| Main Authors: | Srivastava, Sharvani, Singh, Sudhakar, Pooja, Prakash, Shiv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
by: Gan, Yaozong, et al.
Published: (2024)
by: Gan, Yaozong, et al.
Published: (2024)
Testing MediaPipe Holistic for Linguistic Analysis of Nonmanual Markers in Sign Languages
by: Kuznetsova, Anna, et al.
Published: (2024)
by: Kuznetsova, Anna, et al.
Published: (2024)
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
by: Gan, Yaozong, et al.
Published: (2024)
by: Gan, Yaozong, et al.
Published: (2024)
American Sign Language Alphabet Recognition using Deep Learning
by: Kasukurthi, Nikhil, et al.
Published: (2019)
by: Kasukurthi, Nikhil, et al.
Published: (2019)
Hierarchical Sub-action Tree for Continuous Sign Language Recognition
by: Yang, Dejie, et al.
Published: (2025)
by: Yang, Dejie, et al.
Published: (2025)
ForcePose: A Deep Learning Approach for Force Calculation Based on Action Recognition Using MediaPipe Pose Estimation Combined with Object Detection
by: M, Nandakishor, et al.
Published: (2025)
by: M, Nandakishor, et al.
Published: (2025)
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
by: Du, Guanyi, et al.
Published: (2026)
by: Du, Guanyi, et al.
Published: (2026)
Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
by: Yu, Xuzheng, et al.
Published: (2024)
by: Yu, Xuzheng, et al.
Published: (2024)
Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
by: Yu, Zongyou, et al.
Published: (2024)
by: Yu, Zongyou, et al.
Published: (2024)
Attributes-aware Visual Emotion Representation Learning
by: Maharjan, Rahul Singh, et al.
Published: (2025)
by: Maharjan, Rahul Singh, et al.
Published: (2025)
Media Forensics and Deepfake Systematic Survey
by: CH, Nadeem Jabbar, et al.
Published: (2024)
by: CH, Nadeem Jabbar, et al.
Published: (2024)
Efficient Low-Resolution Face Recognition via Bridge Distillation
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
by: Zhang, Junzheng, et al.
Published: (2024)
by: Zhang, Junzheng, et al.
Published: (2024)
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
by: Li, Fuhao, et al.
Published: (2026)
by: Li, Fuhao, et al.
Published: (2026)
Real-Time Posture Monitoring and Risk Assessment for Manual Lifting Tasks Using MediaPipe and LSTM
by: Bagga, Ereena, et al.
Published: (2024)
by: Bagga, Ereena, et al.
Published: (2024)
Automatic Recognition of Food Ingestion Environment from the AIM-2 Wearable Sensor
by: Huang, Yuning, et al.
Published: (2024)
by: Huang, Yuning, et al.
Published: (2024)
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
by: Zhang, Hanwen, et al.
Published: (2026)
by: Zhang, Hanwen, et al.
Published: (2026)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
FakeParts: a New Family of AI-Generated DeepFakes
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images
by: Amoroso, Roberto, et al.
Published: (2023)
by: Amoroso, Roberto, et al.
Published: (2023)
Word-level Sign Language Recognition with Multi-stream Neural Networks Focusing on Local Regions and Skeletal Information
by: Maruyama, Mizuki, et al.
Published: (2021)
by: Maruyama, Mizuki, et al.
Published: (2021)
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Text-Only Data Synthesis for Vision Language Model Training
by: Yu, Xiaomin, et al.
Published: (2025)
by: Yu, Xiaomin, et al.
Published: (2025)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
by: Yu, Lijun, et al.
Published: (2023)
by: Yu, Lijun, et al.
Published: (2023)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
by: Li, Qifei, et al.
Published: (2024)
by: Li, Qifei, et al.
Published: (2024)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
by: Wang, Jiuniu, et al.
Published: (2024)
by: Wang, Jiuniu, et al.
Published: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
by: Lin, Kun-Hsiang, et al.
Published: (2025)
by: Lin, Kun-Hsiang, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
by: Liu, Che, et al.
Published: (2026)
by: Liu, Che, et al.
Published: (2026)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
by: Huang, Feizhen, et al.
Published: (2025)
by: Huang, Feizhen, et al.
Published: (2025)
CountingFruit: Language-Guided 3D Fruit Counting with Semantic Gaussian Splatting
by: Li, Fengze, et al.
Published: (2025)
by: Li, Fengze, et al.
Published: (2025)
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
by: Ouyang, Shuyi, et al.
Published: (2024)
by: Ouyang, Shuyi, et al.
Published: (2024)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
by: Liu, Jiajun, et al.
Published: (2024)
by: Liu, Jiajun, et al.
Published: (2024)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
by: Yu, Xiaomin, et al.
Published: (2026)
by: Yu, Xiaomin, et al.
Published: (2026)
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
by: Zheng, Sixiao, et al.
Published: (2024)
by: Zheng, Sixiao, et al.
Published: (2024)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
by: Xuan, Yunyi, et al.
Published: (2024)
by: Xuan, Yunyi, et al.
Published: (2024)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
by: Ukai, Mahiro, et al.
Published: (2025)
by: Ukai, Mahiro, et al.
Published: (2025)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis
by: Li, Dexiang, et al.
Published: (2026)
by: Li, Dexiang, et al.
Published: (2026)
Similar Items
-
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
by: Gan, Yaozong, et al.
Published: (2024) -
Testing MediaPipe Holistic for Linguistic Analysis of Nonmanual Markers in Sign Languages
by: Kuznetsova, Anna, et al.
Published: (2024) -
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
by: Gan, Yaozong, et al.
Published: (2024) -
American Sign Language Alphabet Recognition using Deep Learning
by: Kasukurthi, Nikhil, et al.
Published: (2019) -
Hierarchical Sub-action Tree for Continuous Sign Language Recognition
by: Yang, Dejie, et al.
Published: (2025)