Gespeichert in:
| Hauptverfasser: | Wang, He, Guo, Pengcheng, Wan, Xucheng, Zhou, Huan, Xie, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.05466 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading
von: Zhang, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhang, Wenhao, et al.
Veröffentlicht: (2024)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
von: Park, Young-Hu, et al.
Veröffentlicht: (2025)
von: Park, Young-Hu, et al.
Veröffentlicht: (2025)
Ev-Layout: A Large-scale Event-based Multi-modal Dataset for Indoor Layout Estimation and Tracking
von: Guo, Xucheng, et al.
Veröffentlicht: (2025)
von: Guo, Xucheng, et al.
Veröffentlicht: (2025)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
von: Huang, Jiehui, et al.
Veröffentlicht: (2025)
von: Huang, Jiehui, et al.
Veröffentlicht: (2025)
VALLR: Visual ASR Language Model for Lip Reading
von: Thomas, Marshall, et al.
Veröffentlicht: (2025)
von: Thomas, Marshall, et al.
Veröffentlicht: (2025)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
von: Wu, Linzhi, et al.
Veröffentlicht: (2024)
von: Wu, Linzhi, et al.
Veröffentlicht: (2024)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
MA-LipNet: Multi-Dimensional Attention Networks for Robust Lipreading
von: Rossi, Matteo
Veröffentlicht: (2026)
von: Rossi, Matteo
Veröffentlicht: (2026)
Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
von: Mao, Junyuan, et al.
Veröffentlicht: (2026)
von: Mao, Junyuan, et al.
Veröffentlicht: (2026)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
von: Maaz, Muhammad, et al.
Veröffentlicht: (2024)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2024)
Hierarchical Semantic Learning for Multi-Class Aorta Segmentation
von: Shi, Pengcheng
Veröffentlicht: (2025)
von: Shi, Pengcheng
Veröffentlicht: (2025)
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
von: He, Zhihao, et al.
Veröffentlicht: (2025)
von: He, Zhihao, et al.
Veröffentlicht: (2025)
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
von: Zhong, Zhizhou, et al.
Veröffentlicht: (2025)
von: Zhong, Zhizhou, et al.
Veröffentlicht: (2025)
YingVideo-MV: Music-Driven Multi-Stage Video Generation
von: Chen, Jiahui, et al.
Veröffentlicht: (2025)
von: Chen, Jiahui, et al.
Veröffentlicht: (2025)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
von: EunGi, Han, et al.
Veröffentlicht: (2024)
von: EunGi, Han, et al.
Veröffentlicht: (2024)
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
von: Peng, Ziqiao, et al.
Veröffentlicht: (2025)
von: Peng, Ziqiao, et al.
Veröffentlicht: (2025)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
von: Huang, Yihuan, et al.
Veröffentlicht: (2025)
von: Huang, Yihuan, et al.
Veröffentlicht: (2025)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
von: Fang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Fang, Pengcheng, et al.
Veröffentlicht: (2025)
LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation
von: Junli, Deng, et al.
Veröffentlicht: (2024)
von: Junli, Deng, et al.
Veröffentlicht: (2024)
Boosting Adversarial Transferability Against Defenses via Multi-Scale Transformation
von: Guo, Zihong, et al.
Veröffentlicht: (2025)
von: Guo, Zihong, et al.
Veröffentlicht: (2025)
Enhance Multi-Scale Spatial-Temporal Coherence for Configurable Video Anomaly Detection
von: Cheng, Kai, et al.
Veröffentlicht: (2023)
von: Cheng, Kai, et al.
Veröffentlicht: (2023)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning
von: He, Yi, et al.
Veröffentlicht: (2025)
von: He, Yi, et al.
Veröffentlicht: (2025)
DynamicLip: Shape-Independent Continuous Authentication via Lip Articulator Dynamics
von: Chen, Huashan, et al.
Veröffentlicht: (2025)
von: Chen, Huashan, et al.
Veröffentlicht: (2025)
Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2024)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2024)
Guided Depth Map Super-Resolution via Multi-Scale Fusion U-shaped Mamba Network
von: Guo, Chenggang, et al.
Veröffentlicht: (2025)
von: Guo, Chenggang, et al.
Veröffentlicht: (2025)
MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
von: Ling, Run, et al.
Veröffentlicht: (2025)
von: Ling, Run, et al.
Veröffentlicht: (2025)
Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
von: Cao, Bing, et al.
Veröffentlicht: (2024)
von: Cao, Bing, et al.
Veröffentlicht: (2024)
Learning Multi-view Anomaly Detection with Efficient Adaptive Selection
von: He, Haoyang, et al.
Veröffentlicht: (2024)
von: He, Haoyang, et al.
Veröffentlicht: (2024)
WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
Integrating Persian Lip Reading in Surena-V Humanoid Robot for Human-Robot Interaction
von: Abbasi, Ali Farshian, et al.
Veröffentlicht: (2025)
von: Abbasi, Ali Farshian, et al.
Veröffentlicht: (2025)
MSVCOD:A Large-Scale Multi-Scene Dataset for Video Camouflage Object Detection
von: Gao, Shuyong, et al.
Veröffentlicht: (2025)
von: Gao, Shuyong, et al.
Veröffentlicht: (2025)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
von: Li, Handong, et al.
Veröffentlicht: (2025)
von: Li, Handong, et al.
Veröffentlicht: (2025)
M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading
von: Zhang, Wenhao, et al.
Veröffentlicht: (2024) -
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
von: Hao, Bowen, et al.
Veröffentlicht: (2025) -
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
von: Park, Young-Hu, et al.
Veröffentlicht: (2025) -
Ev-Layout: A Large-scale Event-based Multi-modal Dataset for Indoor Layout Estimation and Tracking
von: Guo, Xucheng, et al.
Veröffentlicht: (2025) -
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
von: Huang, Jiehui, et al.
Veröffentlicht: (2025)