Text-Driven Diffusion Model for Sign Language Production
Fuente:
arXiv
Saved in:
| Main Authors: | He, Jiayi, Wang, Xu, Zhang, Ruobei, Tang, Shengeng, Wang, Yaxiong, Cheng, Lechao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
by: He, Jiashu, et al.
Published: (2025)
by: He, Jiashu, et al.
Published: (2025)
SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observation
by: Tang, Shengeng, et al.
Published: (2024)
by: Tang, Shengeng, et al.
Published: (2024)
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
by: Tang, Shengeng, et al.
Published: (2024)
by: Tang, Shengeng, et al.
Published: (2024)
Knowledge Swapping via Learning and Unlearning
by: Xing, Mingyu, et al.
Published: (2025)
by: Xing, Mingyu, et al.
Published: (2025)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
by: Wang, Qihang, et al.
Published: (2025)
by: Wang, Qihang, et al.
Published: (2025)
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
by: Li, Haiyang, et al.
Published: (2025)
by: Li, Haiyang, et al.
Published: (2025)
SLRTP2025 Sign Language Production Challenge: Methodology, Results, and Future Work
by: Walsh, Harry, et al.
Published: (2025)
by: Walsh, Harry, et al.
Published: (2025)
Modality Alignment Meets Federated Broadcasting
by: Ma, Yuting, et al.
Published: (2024)
by: Ma, Yuting, et al.
Published: (2024)
Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
by: Guo, Mingce, et al.
Published: (2024)
by: Guo, Mingce, et al.
Published: (2024)
Wi-CBR: Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior Recognition
by: Zhang, Ruobei, et al.
Published: (2025)
by: Zhang, Ruobei, et al.
Published: (2025)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
by: Wang, Xu, et al.
Published: (2024)
by: Wang, Xu, et al.
Published: (2024)
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
by: Shen, Jinjie, et al.
Published: (2026)
by: Shen, Jinjie, et al.
Published: (2026)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
by: Shen, Jinjie, et al.
Published: (2026)
by: Shen, Jinjie, et al.
Published: (2026)
SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition
by: Li, Jiahui, et al.
Published: (2025)
by: Li, Jiahui, et al.
Published: (2025)
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
by: Wang, Yaxiong, et al.
Published: (2025)
by: Wang, Yaxiong, et al.
Published: (2025)
Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning
by: Wu, Fangwen, et al.
Published: (2025)
by: Wu, Fangwen, et al.
Published: (2025)
Dataset Distillers Are Good Label Denoisers In the Wild
by: Cheng, Lechao, et al.
Published: (2024)
by: Cheng, Lechao, et al.
Published: (2024)
SignDiff: Diffusion Model for American Sign Language Production
by: Fang, Sen, et al.
Published: (2023)
by: Fang, Sen, et al.
Published: (2023)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
by: Ma, Yuting, et al.
Published: (2024)
by: Ma, Yuting, et al.
Published: (2024)
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
by: Shen, Jinjie, et al.
Published: (2025)
by: Shen, Jinjie, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
Spoofing-aware Prompt Learning for Unified Physical-Digital Facial Attack Detection
by: Guo, Jiabao, et al.
Published: (2025)
by: Guo, Jiabao, et al.
Published: (2025)
SignLLM: Sign Language Production Large Language Models
by: Fang, Sen, et al.
Published: (2024)
by: Fang, Sen, et al.
Published: (2024)
Hybrid Autoregressive-Diffusion Model for Real-Time Sign Language Production
by: Ye, Maoxiao, et al.
Published: (2025)
by: Ye, Maoxiao, et al.
Published: (2025)
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
by: Zhang, Lechao, et al.
Published: (2026)
by: Zhang, Lechao, et al.
Published: (2026)
Decoupled Training with Local Reinforcement Fine-Tuning in Federated Learning
by: Ma, Yuting, et al.
Published: (2026)
by: Ma, Yuting, et al.
Published: (2026)
Teacher-Student Diffusion Model for Text-Driven 3D Hand Motion Generation
by: Cheng, Ching-Lam, et al.
Published: (2026)
by: Cheng, Ching-Lam, et al.
Published: (2026)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
LoopGaussian: Creating 3D Cinemagraph with Multi-view Images via Eulerian Motion Field
by: Li, Jiyang, et al.
Published: (2024)
by: Li, Jiyang, et al.
Published: (2024)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
by: Wang, Kuo, et al.
Published: (2024)
by: Wang, Kuo, et al.
Published: (2024)
Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Similar Items
-
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
by: He, Jiayi, et al.
Published: (2025) -
CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
by: Wang, Xu, et al.
Published: (2026) -
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
by: He, Jiashu, et al.
Published: (2025) -
SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation
by: Wang, Xu, et al.
Published: (2025) -
Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observation
by: Tang, Shengeng, et al.
Published: (2024)