STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
Fuente:
arXiv
Salvato in:
| Autori principali: | Ding, Zijun, Xiong, Mingdie, Zhu, Congcong, Chen, Jingrun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
di: Liu, Tao, et al.
Pubblicazione: (2023)
di: Liu, Tao, et al.
Pubblicazione: (2023)
Physics-Informed Deformable Gaussian Splatting: Towards Unified Constitutive Laws for Time-Evolving Material Field
di: Hong, Haoqin, et al.
Pubblicazione: (2025)
di: Hong, Haoqin, et al.
Pubblicazione: (2025)
Video Editing for Audio-Visual Dubbing
di: Manela, Binyamin, et al.
Pubblicazione: (2025)
di: Manela, Binyamin, et al.
Pubblicazione: (2025)
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
di: Zhang, Longhao, et al.
Pubblicazione: (2024)
di: Zhang, Longhao, et al.
Pubblicazione: (2024)
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
di: Lin, Ling, et al.
Pubblicazione: (2026)
di: Lin, Ling, et al.
Pubblicazione: (2026)
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding
di: Luo, Bingjun, et al.
Pubblicazione: (2026)
di: Luo, Bingjun, et al.
Pubblicazione: (2026)
Select2Col: Leveraging Spatial-Temporal Importance of Semantic Information for Efficient Collaborative Perception
di: Liu, Yuntao, et al.
Pubblicazione: (2023)
di: Liu, Yuntao, et al.
Pubblicazione: (2023)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
di: Tan, Zhentao, et al.
Pubblicazione: (2024)
di: Tan, Zhentao, et al.
Pubblicazione: (2024)
SSRFlow: Semantic-aware Fusion with Spatial Temporal Re-embedding for Real-world Scene Flow
di: Lu, Zhiyang, et al.
Pubblicazione: (2024)
di: Lu, Zhiyang, et al.
Pubblicazione: (2024)
LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
di: Karimov, Mirlan, et al.
Pubblicazione: (2026)
di: Karimov, Mirlan, et al.
Pubblicazione: (2026)
SwinSF: Image Reconstruction from Spatial-Temporal Spike Streams
di: Jiang, Liangyan, et al.
Pubblicazione: (2024)
di: Jiang, Liangyan, et al.
Pubblicazione: (2024)
ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales
di: Zhen, Yihao, et al.
Pubblicazione: (2025)
di: Zhen, Yihao, et al.
Pubblicazione: (2025)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
di: Nguyen, Ngoc-Son, et al.
Pubblicazione: (2026)
di: Nguyen, Ngoc-Son, et al.
Pubblicazione: (2026)
StarVid: Enhancing Semantic Alignment in Video Diffusion Models via Spatial and SynTactic Guided Attention Refocusing
di: Li, Yuanhang, et al.
Pubblicazione: (2024)
di: Li, Yuanhang, et al.
Pubblicazione: (2024)
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
di: Zheng, Yuze, et al.
Pubblicazione: (2024)
di: Zheng, Yuze, et al.
Pubblicazione: (2024)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
di: Xu, Wenjia, et al.
Pubblicazione: (2024)
di: Xu, Wenjia, et al.
Pubblicazione: (2024)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
di: Sun, Guolei, et al.
Pubblicazione: (2022)
di: Sun, Guolei, et al.
Pubblicazione: (2022)
Non-rigid Structure-from-Motion: Temporally-smooth Procrustean Alignment and Spatially-variant Deformation Modeling
di: Shi, Jiawei, et al.
Pubblicazione: (2024)
di: Shi, Jiawei, et al.
Pubblicazione: (2024)
Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment
di: Bi, Xiaowei, et al.
Pubblicazione: (2025)
di: Bi, Xiaowei, et al.
Pubblicazione: (2025)
Unsupervised Audio-Visual Segmentation with Modality Alignment
di: Bhosale, Swapnil, et al.
Pubblicazione: (2024)
di: Bhosale, Swapnil, et al.
Pubblicazione: (2024)
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
di: Chen, Weiming, et al.
Pubblicazione: (2025)
di: Chen, Weiming, et al.
Pubblicazione: (2025)
Spatial Frequency Modulation for Semantic Segmentation
di: Chen, Linwei, et al.
Pubblicazione: (2025)
di: Chen, Linwei, et al.
Pubblicazione: (2025)
Decoupling Semantic Similarity from Spatial Alignment for Neural Networks
di: Wald, Tassilo, et al.
Pubblicazione: (2024)
di: Wald, Tassilo, et al.
Pubblicazione: (2024)
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
di: Mao, Fangyuan, et al.
Pubblicazione: (2025)
di: Mao, Fangyuan, et al.
Pubblicazione: (2025)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
di: He, Hulingxiao, et al.
Pubblicazione: (2026)
di: He, Hulingxiao, et al.
Pubblicazione: (2026)
Mask-RadarNet: Enhancing Transformer With Spatial-Temporal Semantic Context for Radar Object Detection in Autonomous Driving
di: Wu, Yuzhi, et al.
Pubblicazione: (2024)
di: Wu, Yuzhi, et al.
Pubblicazione: (2024)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
2D Gaussian Splatting with Semantic Alignment for Image Inpainting
di: Li, Hongyu, et al.
Pubblicazione: (2025)
di: Li, Hongyu, et al.
Pubblicazione: (2025)
Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
di: Fan, Senran, et al.
Pubblicazione: (2024)
di: Fan, Senran, et al.
Pubblicazione: (2024)
Beyond Human-prompting: Adaptive Prompt Tuning with Semantic Alignment for Anomaly Detection
di: Chen, Pi-Wei, et al.
Pubblicazione: (2025)
di: Chen, Pi-Wei, et al.
Pubblicazione: (2025)
ESP-PCT: Enhanced VR Semantic Performance through Efficient Compression of Temporal and Spatial Redundancies in Point Cloud Transformers
di: Mei, Luoyu, et al.
Pubblicazione: (2024)
di: Mei, Luoyu, et al.
Pubblicazione: (2024)
FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes
di: Liu, Jiaxuan, et al.
Pubblicazione: (2026)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2026)
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
di: Ju, Shaobo, et al.
Pubblicazione: (2026)
di: Ju, Shaobo, et al.
Pubblicazione: (2026)
Rethinking Alignment and Uniformity in Unsupervised Semantic Segmentation
di: Zhang, Daoan, et al.
Pubblicazione: (2022)
di: Zhang, Daoan, et al.
Pubblicazione: (2022)
Unified Spatial-Temporal Edge-Enhanced Graph Networks for Pedestrian Trajectory Prediction
di: Li, Ruochen, et al.
Pubblicazione: (2025)
di: Li, Ruochen, et al.
Pubblicazione: (2025)
7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
di: Sun, Guoheng, et al.
Pubblicazione: (2026)
di: Sun, Guoheng, et al.
Pubblicazione: (2026)
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
di: Zhu, Zhiyi, et al.
Pubblicazione: (2025)
di: Zhu, Zhiyi, et al.
Pubblicazione: (2025)
Learning Sparse Visual Representations via Spatial-Semantic Factorization
di: Zhao, Theodore Zhengde, et al.
Pubblicazione: (2026)
di: Zhao, Theodore Zhengde, et al.
Pubblicazione: (2026)
Neural Visual Decoding via Cognitive guided Adaptive Blurring and Information Constrained Alignment
di: Yin, Fan, et al.
Pubblicazione: (2026)
di: Yin, Fan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
di: Liu, Tao, et al.
Pubblicazione: (2023) -
Physics-Informed Deformable Gaussian Splatting: Towards Unified Constitutive Laws for Time-Evolving Material Field
di: Hong, Haoqin, et al.
Pubblicazione: (2025) -
Video Editing for Audio-Visual Dubbing
di: Manela, Binyamin, et al.
Pubblicazione: (2025) -
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
di: Zhang, Longhao, et al.
Pubblicazione: (2024) -
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
di: Lin, Ling, et al.
Pubblicazione: (2026)