Scaling up Multimodal Pre-training for Sign Language Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Wengang, Zhao, Weichao, Hu, Hezhen, Li, Zecheng, Li, Houqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uni-Sign: Toward Unified Sign Language Understanding at Scale
von: Li, Zecheng, et al.
Veröffentlicht: (2025)
von: Li, Zecheng, et al.
Veröffentlicht: (2025)
Cross-Modal Consistency Learning for Sign Language Recognition
von: Wu, Kepeng, et al.
Veröffentlicht: (2025)
von: Wu, Kepeng, et al.
Veröffentlicht: (2025)
Self-Supervised Representation Learning with Spatial-Temporal Consistency for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
Exploiting Spatial-Temporal Context for Interacting Hand Reconstruction on Monocular RGB Video
von: Zhao, Weichao, et al.
Veröffentlicht: (2023)
von: Zhao, Weichao, et al.
Veröffentlicht: (2023)
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language Retrieval
von: Jiang, Longtao, et al.
Veröffentlicht: (2024)
von: Jiang, Longtao, et al.
Veröffentlicht: (2024)
Reinforcing Pre-trained Models Using Counterfactual Images
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
Diverse Sign Language Translation
von: Shen, Xin, et al.
Veröffentlicht: (2024)
von: Shen, Xin, et al.
Veröffentlicht: (2024)
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
von: Fu, Jiyuan, et al.
Veröffentlicht: (2024)
von: Fu, Jiyuan, et al.
Veröffentlicht: (2024)
Modularized Zero-shot VQA with Pre-trained Models
von: Cao, Rui, et al.
Veröffentlicht: (2023)
von: Cao, Rui, et al.
Veröffentlicht: (2023)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
Generative Preprocessing for Image Compression with Pre-trained Diffusion Models
von: Guo, Mengxi, et al.
Veröffentlicht: (2025)
von: Guo, Mengxi, et al.
Veröffentlicht: (2025)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
von: Khac, Phúc H. Le, et al.
Veröffentlicht: (2024)
von: Khac, Phúc H. Le, et al.
Veröffentlicht: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Serial Low-rank Adaptation of Vision Transformer
von: Zhong, Houqiang, et al.
Veröffentlicht: (2025)
von: Zhong, Houqiang, et al.
Veröffentlicht: (2025)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
Hierarchical Sub-action Tree for Continuous Sign Language Recognition
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer
von: Luo, Anwei, et al.
Veröffentlicht: (2023)
von: Luo, Anwei, et al.
Veröffentlicht: (2023)
InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
von: Chen, Zhikai, et al.
Veröffentlicht: (2024)
von: Chen, Zhikai, et al.
Veröffentlicht: (2024)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
von: Dong, Xingning, et al.
Veröffentlicht: (2024)
von: Dong, Xingning, et al.
Veröffentlicht: (2024)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
von: Min, Chen, et al.
Veröffentlicht: (2023)
von: Min, Chen, et al.
Veröffentlicht: (2023)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
von: Hao, Jing, et al.
Veröffentlicht: (2025)
von: Hao, Jing, et al.
Veröffentlicht: (2025)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Video-based Sign Language Recognition without Temporal Segmentation
von: Huang, Jie, et al.
Veröffentlicht: (2018)
von: Huang, Jie, et al.
Veröffentlicht: (2018)
Word-level Sign Language Recognition with Multi-stream Neural Networks Focusing on Local Regions and Skeletal Information
von: Maruyama, Mizuki, et al.
Veröffentlicht: (2021)
von: Maruyama, Mizuki, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
Uni-Sign: Toward Unified Sign Language Understanding at Scale
von: Li, Zecheng, et al.
Veröffentlicht: (2025) -
Cross-Modal Consistency Learning for Sign Language Recognition
von: Wu, Kepeng, et al.
Veröffentlicht: (2025) -
Self-Supervised Representation Learning with Spatial-Temporal Consistency for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024) -
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024) -
Exploiting Spatial-Temporal Context for Interacting Hand Reconstruction on Monocular RGB Video
von: Zhao, Weichao, et al.
Veröffentlicht: (2023)