JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Anthony, Korem, Naomi Ken, Zeevi, Gal, Halperin, Tavi, Yosef, Matan Ben, Jelercic, Urska, Bibi, Ofir, Patashnik, Or, Cohen-Or, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVControl: Efficient Framework for Training Audio-Visual Controls
by: Ben-Yosef, Matan, et al.
Published: (2026)
by: Ben-Yosef, Matan, et al.
Published: (2026)
HDR Video Generation via Latent Alignment with Logarithmic Encoding
by: Korem, Naomi Ken, et al.
Published: (2026)
by: Korem, Naomi Ken, et al.
Published: (2026)
TurboEdit: Text-Based Image Editing Using Few-Step Diffusion Models
by: Deutch, Gilad, et al.
Published: (2024)
by: Deutch, Gilad, et al.
Published: (2024)
V-LASIK: Consistent Glasses-Removal from Videos Using Synthetic Data
by: Shalev-Arkushin, Rotem, et al.
Published: (2024)
by: Shalev-Arkushin, Rotem, et al.
Published: (2024)
LCM-Lookahead for Encoder-based Text-to-Image Personalization
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
Dubbing for Everyone: Data-Efficient Visual Dubbing using Neural Rendering Priors
by: Saunders, Jack, et al.
Published: (2024)
by: Saunders, Jack, et al.
Published: (2024)
In-Context Sync-LoRA for Portrait Video Editing
by: Polaczek, Sagi, et al.
Published: (2025)
by: Polaczek, Sagi, et al.
Published: (2025)
LooseRoPE: Content-aware Attention Manipulation for Semantic Harmonization
by: Sella, Etai, et al.
Published: (2026)
by: Sella, Etai, et al.
Published: (2026)
Nested Attention: Semantic-aware Attention Values for Concept Personalization
by: Patashnik, Or, et al.
Published: (2025)
by: Patashnik, Or, et al.
Published: (2025)
Consolidating Attention Features for Multi-view Image Editing
by: Patashnik, Or, et al.
Published: (2024)
by: Patashnik, Or, et al.
Published: (2024)
IP-Composer: Semantic Composition of Visual Concepts
by: Dorfman, Sara, et al.
Published: (2025)
by: Dorfman, Sara, et al.
Published: (2025)
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
by: Zhang, Longhao, et al.
Published: (2024)
by: Zhang, Longhao, et al.
Published: (2024)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
by: Mikaeili, Aryan, et al.
Published: (2026)
by: Mikaeili, Aryan, et al.
Published: (2026)
Be Decisive: Noise-Induced Layouts for Multi-Subject Generation
by: Dahary, Omer, et al.
Published: (2025)
by: Dahary, Omer, et al.
Published: (2025)
Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
by: Dahary, Omer, et al.
Published: (2024)
by: Dahary, Omer, et al.
Published: (2024)
Dynamic Concepts Personalization from Single Videos
by: Abdal, Rameen, et al.
Published: (2025)
by: Abdal, Rameen, et al.
Published: (2025)
Tight Inversion: Image-Conditioned Inversion for Real Image Editing
by: Kadosh, Edo, et al.
Published: (2025)
by: Kadosh, Edo, et al.
Published: (2025)
Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions
by: Sella, Etai, et al.
Published: (2026)
by: Sella, Etai, et al.
Published: (2026)
Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image Generation
by: Cohen, Nadav Z., et al.
Published: (2026)
by: Cohen, Nadav Z., et al.
Published: (2026)
Object-level Visual Prompts for Compositional Image Generation
by: Parmar, Gaurav, et al.
Published: (2025)
by: Parmar, Gaurav, et al.
Published: (2025)
Image Generation from Contextually-Contradictory Prompts
by: Huberman, Saar, et al.
Published: (2025)
by: Huberman, Saar, et al.
Published: (2025)
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
by: Kamenetsky, Ronen, et al.
Published: (2025)
by: Kamenetsky, Ronen, et al.
Published: (2025)
Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation
by: Abramovich, Ofir, et al.
Published: (2026)
by: Abramovich, Ofir, et al.
Published: (2026)
Continuous Control of Editing Models via Adaptive-Origin Guidance
by: Wolf, Alon, et al.
Published: (2026)
by: Wolf, Alon, et al.
Published: (2026)
VLM-Guided Adaptive Negative Prompting for Creative Generation
by: Golan, Shelly, et al.
Published: (2025)
by: Golan, Shelly, et al.
Published: (2025)
S‐ACORD: Spectral Analysis of COral Reef Deformation
by: Naama Alon‐Borissiouk, et al.
Published: (2025)
by: Naama Alon‐Borissiouk, et al.
Published: (2025)
TriTex: Learning Texture from a Single Mesh via Triplane Semantic Features
by: Cohen-Bar, Dana, et al.
Published: (2025)
by: Cohen-Bar, Dana, et al.
Published: (2025)
Stable Flow: Vital Layers for Training-Free Image Editing
by: Avrahami, Omri, et al.
Published: (2024)
by: Avrahami, Omri, et al.
Published: (2024)
ReNoise: Real Image Inversion Through Iterative Noising
by: Garibi, Daniel, et al.
Published: (2024)
by: Garibi, Daniel, et al.
Published: (2024)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
4‐LEGS: 4D Language Embedded Gaussian Splatting
by: Gal Fiebelman, et al.
Published: (2025)
by: Gal Fiebelman, et al.
Published: (2025)
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model
by: Mu, Lingzhou, et al.
Published: (2025)
by: Mu, Lingzhou, et al.
Published: (2025)
MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance
by: Sala, Nathan, et al.
Published: (2026)
by: Sala, Nathan, et al.
Published: (2026)
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
by: Abdal, Rameen, et al.
Published: (2025)
by: Abdal, Rameen, et al.
Published: (2025)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
4-LEGS: 4D Language Embedded Gaussian Splatting
by: Fiebelman, Gal, et al.
Published: (2024)
by: Fiebelman, Gal, et al.
Published: (2024)
Alterbute: Editing Intrinsic Attributes of Objects in Images
by: Reiss, Tal, et al.
Published: (2026)
by: Reiss, Tal, et al.
Published: (2026)
Creative Blends of Visual Concepts
by: Sun, Zhida, et al.
Published: (2025)
by: Sun, Zhida, et al.
Published: (2025)
PlaMo: Plan and Move in Rich 3D Physical Environments
by: Hallak, Assaf, et al.
Published: (2024)
by: Hallak, Assaf, et al.
Published: (2024)
VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models
by: Kim, Geonung, et al.
Published: (2025)
by: Kim, Geonung, et al.
Published: (2025)
Similar Items
-
AVControl: Efficient Framework for Training Audio-Visual Controls
by: Ben-Yosef, Matan, et al.
Published: (2026) -
HDR Video Generation via Latent Alignment with Logarithmic Encoding
by: Korem, Naomi Ken, et al.
Published: (2026) -
TurboEdit: Text-Based Image Editing Using Few-Step Diffusion Models
by: Deutch, Gilad, et al.
Published: (2024) -
V-LASIK: Consistent Glasses-Removal from Videos Using Synthetic Data
by: Shalev-Arkushin, Rotem, et al.
Published: (2024) -
LCM-Lookahead for Encoder-based Text-to-Image Personalization
by: Gal, Rinon, et al.
Published: (2024)