Video Editing for Audio-Visual Dubbing
Fuente:
arXiv
Saved in:
| Main Authors: | Manela, Binyamin, Gannot, Sharon, Fetyaya, Ethan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DAVE: Diagnostic benchmark for Audio Visual Evaluation
by: Radevski, Gorjan, et al.
Published: (2025)
by: Radevski, Gorjan, et al.
Published: (2025)
Pathways on the Image Manifold: Image Editing via Video Generation
by: Rotstein, Noam, et al.
Published: (2024)
by: Rotstein, Noam, et al.
Published: (2024)
Automated Detection of Sport Highlights from Audio and Video Sources
by: Della Santa, Francesco, et al.
Published: (2025)
by: Della Santa, Francesco, et al.
Published: (2025)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
by: Jeong, Hyeonho, et al.
Published: (2023)
by: Jeong, Hyeonho, et al.
Published: (2023)
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
by: Lee, Dohun, et al.
Published: (2026)
by: Lee, Dohun, et al.
Published: (2026)
Seeing Voices: Generating A-Roll Video from Audio with Mirage
by: Sundararaman, Aditi, et al.
Published: (2025)
by: Sundararaman, Aditi, et al.
Published: (2025)
DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
by: Liu, Tao, et al.
Published: (2023)
by: Liu, Tao, et al.
Published: (2023)
X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
by: Kim, Youngseo, et al.
Published: (2026)
by: Kim, Youngseo, et al.
Published: (2026)
VideoNSA: Native Sparse Attention Scales Video Understanding
by: Song, Enxin, et al.
Published: (2025)
by: Song, Enxin, et al.
Published: (2025)
When Test-Time Guidance Is Enough: Fast Image and Video Editing with Diffusion Guidance
by: Ghorbel, Ahmed, et al.
Published: (2026)
by: Ghorbel, Ahmed, et al.
Published: (2026)
Diffusion Model-Based Video Editing: A Survey
by: Sun, Wenhao, et al.
Published: (2024)
by: Sun, Wenhao, et al.
Published: (2024)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
by: Cao-Dinh, Duc, et al.
Published: (2025)
by: Cao-Dinh, Duc, et al.
Published: (2025)
MotionV2V: Editing Motion in a Video
by: Burgert, Ryan, et al.
Published: (2025)
by: Burgert, Ryan, et al.
Published: (2025)
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
by: Ding, Zijun, et al.
Published: (2025)
by: Ding, Zijun, et al.
Published: (2025)
Hearing Touch: Audio-Visual Pretraining for Contact-Rich Manipulation
by: Mejia, Jared, et al.
Published: (2024)
by: Mejia, Jared, et al.
Published: (2024)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
Visual Encoders for Data-Efficient Imitation Learning in Modern Video Games
by: Schäfer, Lukas, et al.
Published: (2023)
by: Schäfer, Lukas, et al.
Published: (2023)
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
by: Croitoru, Florinel-Alin, et al.
Published: (2025)
by: Croitoru, Florinel-Alin, et al.
Published: (2025)
Enhancing Autonomous Vehicle Perception in Adverse Weather through Image Augmentation during Semantic Segmentation Training
by: Kou, Ethan, et al.
Published: (2024)
by: Kou, Ethan, et al.
Published: (2024)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
VINCIE: Unlocking In-context Image Editing from Video
by: Qu, Leigang, et al.
Published: (2025)
by: Qu, Leigang, et al.
Published: (2025)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
by: Barreto, Jesimon, et al.
Published: (2025)
by: Barreto, Jesimon, et al.
Published: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
Uncovering and Mitigating Transient Blindness in Multimodal Model Editing
by: Han, Xiaoqi, et al.
Published: (2025)
by: Han, Xiaoqi, et al.
Published: (2025)
Text-Driven Image Editing via Learnable Regions
by: Lin, Yuanze, et al.
Published: (2023)
by: Lin, Yuanze, et al.
Published: (2023)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
ToDo: Token Downsampling for Efficient Generation of High-Resolution Images
by: Smith, Ethan, et al.
Published: (2024)
by: Smith, Ethan, et al.
Published: (2024)
Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
by: Yao, David Yifan, et al.
Published: (2025)
by: Yao, David Yifan, et al.
Published: (2025)
3D-Consistent Multi-View Editing by Correspondence Guidance
by: Bengtson, Josef, et al.
Published: (2025)
by: Bengtson, Josef, et al.
Published: (2025)
Uniform Attention Maps: Boosting Image Fidelity in Reconstruction and Editing
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
DSCA: Dynamic Subspace Concept Alignment for Lifelong VLM Editing
by: Das, Gyanendra, et al.
Published: (2026)
by: Das, Gyanendra, et al.
Published: (2026)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
by: Iakovleva, Ekaterina, et al.
Published: (2024)
by: Iakovleva, Ekaterina, et al.
Published: (2024)
SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Improving Diffusion-Based Image Editing Faithfulness via Guidance and Scheduling
by: Cho, Hansam, et al.
Published: (2025)
by: Cho, Hansam, et al.
Published: (2025)
SynVA: A Modular Toolkit for Vessel Generation and Aneurysm Editing
by: Finck, Marten J., et al.
Published: (2026)
by: Finck, Marten J., et al.
Published: (2026)
Similar Items
-
DAVE: Diagnostic benchmark for Audio Visual Evaluation
by: Radevski, Gorjan, et al.
Published: (2025) -
Pathways on the Image Manifold: Image Editing via Video Generation
by: Rotstein, Noam, et al.
Published: (2024) -
Automated Detection of Sport Highlights from Audio and Video Sources
by: Della Santa, Francesco, et al.
Published: (2025) -
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024) -
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
by: Jeong, Hyeonho, et al.
Published: (2023)