VALLR-Pin: Uncertainty-Factorized Visual Speech Recognition for Mandarin with Pinyin Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Chang, Xie, Dongliang, Xie, Wanpeng, Qin, Bo, Yang, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VALLR: Visual ASR Language Model for Lip Reading
by: Thomas, Marshall, et al.
Published: (2025)
by: Thomas, Marshall, et al.
Published: (2025)
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
by: Yang, Lei, et al.
Published: (2026)
by: Yang, Lei, et al.
Published: (2026)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025)
by: Hao, Bowen, et al.
Published: (2025)
The NPU-ASLP System Description for Visual Speech Recognition in CNVSRC 2024
by: Wang, He, et al.
Published: (2024)
by: Wang, He, et al.
Published: (2024)
LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
by: Wang, Hefeng, et al.
Published: (2024)
by: Wang, Hefeng, et al.
Published: (2024)
Dynamic Resolution Guidance for Facial Expression Recognition
by: Wang, Songpan, et al.
Published: (2024)
by: Wang, Songpan, et al.
Published: (2024)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
Streamlined Open-Vocabulary Human-Object Interaction Detection
by: Sun, Chang, et al.
Published: (2026)
by: Sun, Chang, et al.
Published: (2026)
SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
by: Xu, Dongli, et al.
Published: (2025)
by: Xu, Dongli, et al.
Published: (2025)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
by: Wu, Junfeng, et al.
Published: (2025)
by: Wu, Junfeng, et al.
Published: (2025)
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
by: Xie, Peijin, et al.
Published: (2025)
by: Xie, Peijin, et al.
Published: (2025)
Towards Privacy-Preserving Fine-Grained Visual Classification via Hierarchical Learning from Label Proportions
by: Chang, Jinyi, et al.
Published: (2025)
by: Chang, Jinyi, et al.
Published: (2025)
Multi-view Normal and Distance Guidance Gaussian Splatting for Surface Reconstruction
by: Jia, Bo, et al.
Published: (2025)
by: Jia, Bo, et al.
Published: (2025)
On the Estimation of Image-matching Uncertainty in Visual Place Recognition
by: Zaffar, Mubariz, et al.
Published: (2024)
by: Zaffar, Mubariz, et al.
Published: (2024)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
by: Li, Cancan, et al.
Published: (2025)
by: Li, Cancan, et al.
Published: (2025)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance
by: Pan, Dongwei, et al.
Published: (2026)
by: Pan, Dongwei, et al.
Published: (2026)
Mamba-Adaptor: State Space Model Adaptor for Visual Recognition
by: Xie, Fei, et al.
Published: (2025)
by: Xie, Fei, et al.
Published: (2025)
GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition
by: Wang, Tianyue, et al.
Published: (2025)
by: Wang, Tianyue, et al.
Published: (2025)
Transfer Learning from Visual Speech Recognition to Mouthing Recognition in German Sign Language
by: Pham, Dinh Nam, et al.
Published: (2025)
by: Pham, Dinh Nam, et al.
Published: (2025)
PIG: Prompt Images Guidance for Night-Time Scene Parsing
by: Xie, Zhifeng, et al.
Published: (2024)
by: Xie, Zhifeng, et al.
Published: (2024)
Enhanced Semantic Extraction and Guidance for UGC Image Super Resolution
by: Wang, Yiwen, et al.
Published: (2025)
by: Wang, Yiwen, et al.
Published: (2025)
Online Navigation Refinement: Achieving Lane-Level Guidance by Associating Standard-Definition and Online Perception Maps
by: Wan, Jiaxu, et al.
Published: (2025)
by: Wan, Jiaxu, et al.
Published: (2025)
Federated Dialogue-Semantic Diffusion for Emotion Recognition under Incomplete Modalities
by: Qiu, Xihang, et al.
Published: (2025)
by: Qiu, Xihang, et al.
Published: (2025)
Robust 6DoF Pose Tracking Considering Contour and Interior Correspondence Uncertainty for AR Assembly Guidance
by: Chen, Jixiang, et al.
Published: (2025)
by: Chen, Jixiang, et al.
Published: (2025)
Head-Pose-Aware Visual Speech Recognition with FiLM Modulation
by: Teng, Matthew Kit Khinn, et al.
Published: (2026)
by: Teng, Matthew Kit Khinn, et al.
Published: (2026)
Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector
by: Zu, Wenqiang, et al.
Published: (2026)
by: Zu, Wenqiang, et al.
Published: (2026)
Visual Prompting in LLMs for Enhancing Emotion Recognition
by: Zhang, Qixuan, et al.
Published: (2024)
by: Zhang, Qixuan, et al.
Published: (2024)
JoyHallo: Digital human model for Mandarin
by: Shi, Sheng, et al.
Published: (2024)
by: Shi, Sheng, et al.
Published: (2024)
Visual-Geometric Collaborative Guidance for Affordance Learning
by: Luo, Hongchen, et al.
Published: (2024)
by: Luo, Hongchen, et al.
Published: (2024)
How is Visual Attention Influenced by Text Guidance? Database and Model
by: Sun, Yinan, et al.
Published: (2024)
by: Sun, Yinan, et al.
Published: (2024)
StructInbet: Integrating Explicit Structural Guidance into Inbetween Frame Generation
by: Pan, Zhenglin, et al.
Published: (2025)
by: Pan, Zhenglin, et al.
Published: (2025)
BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition
by: Gimeno-Gómez, David, et al.
Published: (2024)
by: Gimeno-Gómez, David, et al.
Published: (2024)
Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
by: Chen, Junhan, et al.
Published: (2026)
by: Chen, Junhan, et al.
Published: (2026)
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
Similar Items
-
VALLR: Visual ASR Language Model for Lip Reading
by: Thomas, Marshall, et al.
Published: (2025) -
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition
by: Sun, Chang, et al.
Published: (2024) -
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
by: Yang, Lei, et al.
Published: (2026) -
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
by: Hao, Bowen, et al.
Published: (2025) -
The NPU-ASLP System Description for Visual Speech Recognition in CNVSRC 2024
by: Wang, He, et al.
Published: (2024)