Hierarchical Sub-action Tree for Continuous Sign Language Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Dejie, Xu, Zhu, Gao, Xinjie, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
Word-level Sign Language Recognition with Multi-stream Neural Networks Focusing on Local Regions and Skeletal Information
von: Maruyama, Mizuki, et al.
Veröffentlicht: (2021)
von: Maruyama, Mizuki, et al.
Veröffentlicht: (2021)
Diverse Sign Language Translation
von: Shen, Xin, et al.
Veröffentlicht: (2024)
von: Shen, Xin, et al.
Veröffentlicht: (2024)
A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
von: He, Zongtao, et al.
Veröffentlicht: (2023)
von: He, Zongtao, et al.
Veröffentlicht: (2023)
Learning Brain Representation with Hierarchical Visual Embeddings
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
Scaling up Multimodal Pre-training for Sign Language Understanding
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
Continuous Sign Language Recognition System using Deep Learning with MediaPipe Holistic
von: Srivastava, Sharvani, et al.
Veröffentlicht: (2024)
von: Srivastava, Sharvani, et al.
Veröffentlicht: (2024)
Continuous Patch Stitching for Block-wise Image Compression
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
von: Liu, Ke, et al.
Veröffentlicht: (2025)
von: Liu, Ke, et al.
Veröffentlicht: (2025)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
WDMIR: Wavelet-Driven Multimodal Intent Recognition
von: Gong, Weiyin, et al.
Veröffentlicht: (2025)
von: Gong, Weiyin, et al.
Veröffentlicht: (2025)
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation
von: Lyu, Guangtao, et al.
Veröffentlicht: (2025)
von: Lyu, Guangtao, et al.
Veröffentlicht: (2025)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Dynamic Resolution Guidance for Facial Expression Recognition
von: Wang, Songpan, et al.
Veröffentlicht: (2024)
von: Wang, Songpan, et al.
Veröffentlicht: (2024)
Improving Gloss-free Sign Language Translation by Reducing Representation Density
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
von: Zhou, Shengli, et al.
Veröffentlicht: (2025)
von: Zhou, Shengli, et al.
Veröffentlicht: (2025)
DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
von: Yang, Wei, et al.
Veröffentlicht: (2025)
von: Yang, Wei, et al.
Veröffentlicht: (2025)
Textured mesh Quality Assessment using Geometry and Color Field Similarity
von: Yang, Kaifa, et al.
Veröffentlicht: (2025)
von: Yang, Kaifa, et al.
Veröffentlicht: (2025)
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
von: Dong, Wenqi, et al.
Veröffentlicht: (2025)
von: Dong, Wenqi, et al.
Veröffentlicht: (2025)
Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
InstructHumans: Editing Animated 3D Human Textures with Instructions
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024) -
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
von: Zhang, Rui, et al.
Veröffentlicht: (2024) -
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
von: Tang, Shengeng, et al.
Veröffentlicht: (2024) -
Word-level Sign Language Recognition with Multi-stream Neural Networks Focusing on Local Regions and Skeletal Information
von: Maruyama, Mizuki, et al.
Veröffentlicht: (2021) -
Diverse Sign Language Translation
von: Shen, Xin, et al.
Veröffentlicht: (2024)