POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Joshi, Abhinav, Sharma, Vaibhav, Singh, Sanjeet, Modi, Ashutosh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
iSign: A Benchmark for Indian Sign Language Processing
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation
by: Tu, Guobin, et al.
Published: (2026)
by: Tu, Guobin, et al.
Published: (2026)
AutoSign: Direct Pose-to-Text Translation for Continuous Sign Language Recognition
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages
by: Fang, Sen, et al.
Published: (2026)
by: Fang, Sen, et al.
Published: (2026)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
by: Zheng, Qiaoyu, et al.
Published: (2025)
by: Zheng, Qiaoyu, et al.
Published: (2025)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
by: Ahn, Young Jin, et al.
Published: (2024)
by: Ahn, Young Jin, et al.
Published: (2024)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
by: Xie, Xudong, et al.
Published: (2024)
by: Xie, Xudong, et al.
Published: (2024)
SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale
by: Gueuwou, Shester, et al.
Published: (2024)
by: Gueuwou, Shester, et al.
Published: (2024)
Generation and Detection of Sign Language Deepfakes - A Linguistic and Visual Analysis
by: Naeem, Shahzeb, et al.
Published: (2024)
by: Naeem, Shahzeb, et al.
Published: (2024)
Towards Privacy-Aware Sign Language Translation at Scale
by: Rust, Phillip, et al.
Published: (2024)
by: Rust, Phillip, et al.
Published: (2024)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
by: Zhao, Zhonghan, et al.
Published: (2025)
by: Zhao, Zhonghan, et al.
Published: (2025)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
by: Mohamed, Ibrahim Sheikh, et al.
Published: (2025)
by: Mohamed, Ibrahim Sheikh, et al.
Published: (2025)
Sign Stitching: A Novel Approach to Sign Language Production
by: Walsh, Harry, et al.
Published: (2024)
by: Walsh, Harry, et al.
Published: (2024)
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
by: Khalil, Mohammad Amer, et al.
Published: (2026)
by: Khalil, Mohammad Amer, et al.
Published: (2026)
Augmenting End-to-End Steering Angle Prediction with CAN Bus Data
by: Singh, Amit
Published: (2023)
by: Singh, Amit
Published: (2023)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
by: Liang, Han, et al.
Published: (2024)
by: Liang, Han, et al.
Published: (2024)
EMMA: End-to-End Multimodal Model for Autonomous Driving
by: Hwang, Jyh-Jing, et al.
Published: (2024)
by: Hwang, Jyh-Jing, et al.
Published: (2024)
3D-LEX v1.0: 3D Lexicons for American Sign Language and Sign Language of the Netherlands
by: Ranum, Oline, et al.
Published: (2024)
by: Ranum, Oline, et al.
Published: (2024)
GutenOCR: A Grounded Vision-Language Front-End for Documents
by: Heidenreich, Hunter, et al.
Published: (2026)
by: Heidenreich, Hunter, et al.
Published: (2026)
Beyond Vision: How Large Language Models Interpret Facial Expressions from Valence-Arousal Values
by: Mehra, Vaibhav, et al.
Published: (2025)
by: Mehra, Vaibhav, et al.
Published: (2025)
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
by: Zhang, Yaping, et al.
Published: (2026)
by: Zhang, Yaping, et al.
Published: (2026)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
by: Elobaid, Alaa
Published: (2026)
by: Elobaid, Alaa
Published: (2026)
Geometry of Decision Making in Language Models
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model
by: Wu, Zhanglin, et al.
Published: (2025)
by: Wu, Zhanglin, et al.
Published: (2025)
Modeling Intensification for Sign Language Generation: A Computational Approach
by: İnan, Mert, et al.
Published: (2022)
by: İnan, Mert, et al.
Published: (2022)
SCOPE: Sign Language Contextual Processing with Embedding from LLMs
by: Liu, Yuqi, et al.
Published: (2024)
by: Liu, Yuqi, et al.
Published: (2024)
DiffSLT: Enhancing Diversity in Sign Language Translation via Diffusion Model
by: Moon, JiHwan, et al.
Published: (2024)
by: Moon, JiHwan, et al.
Published: (2024)
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
by: Gao, Yufei, et al.
Published: (2025)
by: Gao, Yufei, et al.
Published: (2025)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
by: Singh, Anshul, et al.
Published: (2025)
by: Singh, Anshul, et al.
Published: (2025)
A Comprehensive Review of Sign Language Recognition: Different Types, Modalities, and Datasets
by: Madhiarasan, M., et al.
Published: (2022)
by: Madhiarasan, M., et al.
Published: (2022)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
by: Deng, Ken, et al.
Published: (2026)
by: Deng, Ken, et al.
Published: (2026)
Event Stream-based Sign Language Translation: A High-Definition Benchmark Dataset and A Novel Baseline
by: Wang, Shiao, et al.
Published: (2024)
by: Wang, Shiao, et al.
Published: (2024)
End-to-End Human Instance Matting
by: Liu, Qinglin, et al.
Published: (2024)
by: Liu, Qinglin, et al.
Published: (2024)
Show Me the World in My Language: Establishing the First Baseline for Scene-Text to Scene-Text Translation
by: Vaidya, Shreyas, et al.
Published: (2023)
by: Vaidya, Shreyas, et al.
Published: (2023)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
by: Jiang, Hao, et al.
Published: (2025)
by: Jiang, Hao, et al.
Published: (2025)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
by: Zhu, Wang, et al.
Published: (2023)
by: Zhu, Wang, et al.
Published: (2023)
Similar Items
-
iSign: A Benchmark for Indian Sign Language Processing
by: Joshi, Abhinav, et al.
Published: (2024) -
FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation
by: Tu, Guobin, et al.
Published: (2026) -
AutoSign: Direct Pose-to-Text Translation for Continuous Sign Language Recognition
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025) -
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025) -
SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages
by: Fang, Sen, et al.
Published: (2026)