InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Xue, Junxiao, Liu, Xiaozhen, Wu, Xuecheng, Yu, Fei, Wang, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
A Trustworthy Method for Multimodal Emotion Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition
by: Cui, Kai, et al.
Published: (2025)
by: Cui, Kai, et al.
Published: (2025)
UniSync: A Unified Framework for Audio-Visual Synchronization
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
Affective Video Content Analysis: Decade Review and New Perspectives
by: Xue, Junxiao, et al.
Published: (2023)
by: Xue, Junxiao, et al.
Published: (2023)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
by: Ahn, Young Jin, et al.
Published: (2024)
by: Ahn, Young Jin, et al.
Published: (2024)
Interpretable Convolutional SyncNet
by: Park, Sungjoon, et al.
Published: (2024)
by: Park, Sungjoon, et al.
Published: (2024)
SyncVIS: Synchronized Video Instance Segmentation
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024)
by: Xue, Junxiao, et al.
Published: (2024)
RocSync: Millisecond-Accurate Temporal Synchronization for Heterogeneous Camera Systems
by: Meyer, Jaro, et al.
Published: (2025)
by: Meyer, Jaro, et al.
Published: (2025)
SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
by: Peng, Ziqiao, et al.
Published: (2023)
by: Peng, Ziqiao, et al.
Published: (2023)
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
by: Xue, Junxiao, et al.
Published: (2026)
by: Xue, Junxiao, et al.
Published: (2026)
SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning
by: Cheng, Xin, et al.
Published: (2026)
by: Cheng, Xin, et al.
Published: (2026)
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
KAConvNet: Kolmogorov-Arnold Convolutional Networks for Vision Recognition
by: Liu, Zhaoxiang, et al.
Published: (2026)
by: Liu, Zhaoxiang, et al.
Published: (2026)
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
UniSync: Towards Generalizable and High-Fidelity Lip Synchronization for Challenging Scenarios
by: Fan, Ruidi, et al.
Published: (2026)
by: Fan, Ruidi, et al.
Published: (2026)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
by: Huang, Yihuan, et al.
Published: (2025)
by: Huang, Yihuan, et al.
Published: (2025)
3A-YOLO: New Real-Time Object Detectors with Triple Discriminative Awareness and Coordinated Representations
by: Wu, Xuecheng, et al.
Published: (2024)
by: Wu, Xuecheng, et al.
Published: (2024)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
by: Li, Chunyu, et al.
Published: (2024)
by: Li, Chunyu, et al.
Published: (2024)
Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes
by: Liu, Weifeng, et al.
Published: (2024)
by: Liu, Weifeng, et al.
Published: (2024)
eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
by: Liu, Shaowei, et al.
Published: (2025)
by: Liu, Shaowei, et al.
Published: (2025)
Disentangling Hardness from Noise: An Uncertainty-Driven Model-Agnostic Framework for Long-Tailed Remote Sensing Classification
by: Ding, Chi, et al.
Published: (2026)
by: Ding, Chi, et al.
Published: (2026)
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
Edge-Cloud Collaborative Satellite Image Analysis for Efficient Man-Made Structure Recognition
by: Sheng, Kaicheng, et al.
Published: (2024)
by: Sheng, Kaicheng, et al.
Published: (2024)
Ghost-dil-NetVLAD: A Lightweight Neural Network for Visual Place Recognition
by: Gong, Qingyuan, et al.
Published: (2021)
by: Gong, Qingyuan, et al.
Published: (2021)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs
by: Zhou, Ziheng, et al.
Published: (2026)
by: Zhou, Ziheng, et al.
Published: (2026)
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
by: Yang, Lei, et al.
Published: (2026)
by: Yang, Lei, et al.
Published: (2026)
SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
Local Spatiotemporal Convolutional Network for Robust Gait Recognition
by: Wang, Xiaoyun, et al.
Published: (2026)
by: Wang, Xiaoyun, et al.
Published: (2026)
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
SyncTweedies: A General Generative Framework Based on Synchronized Diffusions
by: Kim, Jaihoon, et al.
Published: (2024)
by: Kim, Jaihoon, et al.
Published: (2024)
Boosting Continuous Emotion Recognition with Self-Pretraining using Masked Autoencoders, Temporal Convolutional Networks, and Transformers
by: Zhou, Weiwei, et al.
Published: (2024)
by: Zhou, Weiwei, et al.
Published: (2024)
Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline
by: Wu, Xuecheng, et al.
Published: (2023)
by: Wu, Xuecheng, et al.
Published: (2023)
HighSync: High-Quality Lip Synchronization via Latent Diffusion Models
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization
by: Li, Deming, et al.
Published: (2026)
by: Li, Deming, et al.
Published: (2026)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
by: Liang, Sen, et al.
Published: (2026)
by: Liang, Sen, et al.
Published: (2026)
Similar Items
-
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025) -
A Trustworthy Method for Multimodal Emotion Recognition
by: Xue, Junxiao, et al.
Published: (2025) -
PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition
by: Cui, Kai, et al.
Published: (2025) -
UniSync: A Unified Framework for Audio-Visual Synchronization
by: Feng, Tao, et al.
Published: (2025) -
Affective Video Content Analysis: Decade Review and New Perspectives
by: Xue, Junxiao, et al.
Published: (2023)