Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Rui, Kosecka, Jana |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
by: Rajabi, Navid, et al.
Published: (2023)
by: Rajabi, Navid, et al.
Published: (2023)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
PointSplat: Efficient Geometry-Driven Pruning and Transformer Refinement for 3D Gaussian Splatting
by: Tran, Anh Thuan, et al.
Published: (2026)
by: Tran, Anh Thuan, et al.
Published: (2026)
VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAM
by: Tran, Anh Thuan, et al.
Published: (2026)
by: Tran, Anh Thuan, et al.
Published: (2026)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
by: Rajabi, Navid, et al.
Published: (2025)
by: Rajabi, Navid, et al.
Published: (2025)
Structured Spatial Reasoning with Open Vocabulary Object Detectors
by: Nejatishahidin, Negar, et al.
Published: (2024)
by: Nejatishahidin, Negar, et al.
Published: (2024)
Compositional Image-Text Matching and Retrieval by Grounding Entities
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
by: He, Jiashu, et al.
Published: (2025)
by: He, Jiashu, et al.
Published: (2025)
Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
by: Kezar, Lee, et al.
Published: (2025)
by: Kezar, Lee, et al.
Published: (2025)
Multi-temporal Adaptive Red-Green-Blue and Long-Wave Infrared Fusion for You Only Look Once-Based Landmine Detection from Unmanned Aerial Systems
by: Gallagher, James E., et al.
Published: (2025)
by: Gallagher, James E., et al.
Published: (2025)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)
by: Beňová, Ivana, et al.
Published: (2024)
SignAvatar: Sign Language 3D Motion Reconstruction and Generation
by: Dong, Lu, et al.
Published: (2024)
by: Dong, Lu, et al.
Published: (2024)
Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition
by: Cheng, Bryan, et al.
Published: (2026)
by: Cheng, Bryan, et al.
Published: (2026)
WLASL-LEX: a Dataset for Recognising Phonological Properties in American Sign Language
by: Tavella, Federico, et al.
Published: (2022)
by: Tavella, Federico, et al.
Published: (2022)
A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization
by: Liu, Hanchao, et al.
Published: (2026)
by: Liu, Hanchao, et al.
Published: (2026)
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
by: Tang, Shengeng, et al.
Published: (2024)
by: Tang, Shengeng, et al.
Published: (2024)
GeoDiffMM: Geometry-Guided Conditional Diffusion for Motion Magnification
by: Liu, Xuedeng, et al.
Published: (2025)
by: Liu, Xuedeng, et al.
Published: (2025)
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
by: Hou, Ruibing, et al.
Published: (2026)
by: Hou, Ruibing, et al.
Published: (2026)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Pose-Guided Fine-Grained Sign Language Video Generation
by: Shi, Tongkai, et al.
Published: (2024)
by: Shi, Tongkai, et al.
Published: (2024)
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
SignDiff: Diffusion Model for American Sign Language Production
by: Fang, Sen, et al.
Published: (2023)
by: Fang, Sen, et al.
Published: (2023)
Modelling the Distribution of Human Motion for Sign Language Assessment
by: Cory, Oliver, et al.
Published: (2024)
by: Cory, Oliver, et al.
Published: (2024)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026)
by: Zuo, Ronglai, et al.
Published: (2026)
Towards Continuous Sign Language Conversation from Isolated Signs
by: Kim, Youngmin, et al.
Published: (2026)
by: Kim, Youngmin, et al.
Published: (2026)
Uni-Sign: Toward Unified Sign Language Understanding at Scale
by: Li, Zecheng, et al.
Published: (2025)
by: Li, Zecheng, et al.
Published: (2025)
Shape Conditioned Human Motion Generation with Diffusion Model
by: Xue, Kebing, et al.
Published: (2024)
by: Xue, Kebing, et al.
Published: (2024)
SignAvatars: A Large-scale 3D Sign Language Holistic Motion Dataset and Benchmark
by: Yu, Zhengdi, et al.
Published: (2023)
by: Yu, Zhengdi, et al.
Published: (2023)
Bilingual Text-to-Motion Generation: A New Benchmark and Baselines
by: Weng, Wanjiang, et al.
Published: (2026)
by: Weng, Wanjiang, et al.
Published: (2026)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
by: Hu, Yaosi, et al.
Published: (2023)
by: Hu, Yaosi, et al.
Published: (2023)
Language-Guided Transformer Tokenizer for Human Motion Generation
by: Yan, Sheng, et al.
Published: (2026)
by: Yan, Sheng, et al.
Published: (2026)
EASL: Multi-Emotion Guided Semantic Disentanglement for Expressive Sign Language Generation
by: Zhao, Yanchao, et al.
Published: (2025)
by: Zhao, Yanchao, et al.
Published: (2025)
GLOS: Sign Language Generation with Temporally Aligned Gloss-Level Conditioning
by: Lee, Taeryung, et al.
Published: (2025)
by: Lee, Taeryung, et al.
Published: (2025)
Similar Items
-
Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument
by: Hong, Rui, et al.
Published: (2026) -
Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
by: Hong, Rui, et al.
Published: (2026) -
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
by: Rajabi, Navid, et al.
Published: (2023) -
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
by: Fayyazsanavi, Pooya, et al.
Published: (2024) -
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
by: Rajabi, Navid, et al.
Published: (2024)