SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Sen, Wang, Cong, Guan, Fengbin, Yu, Zhentao, Lu, Yiting, Wang, Yuanzhi, Zhou, Yuan, Li, Xin, Chen, Zhibo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video Quality Assessment Based on Swin TransformerV2 and Coarse to Fine Strategy
by: Yu, Zihao, et al.
Published: (2024)
by: Yu, Zihao, et al.
Published: (2024)
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model
by: Guan, Fengbin, et al.
Published: (2025)
by: Guan, Fengbin, et al.
Published: (2025)
QMamba: On First Exploration of Vision Mamba for Image Quality Assessment
by: Guan, Fengbin, et al.
Published: (2024)
by: Guan, Fengbin, et al.
Published: (2024)
Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
by: Liang, Sen, et al.
Published: (2025)
by: Liang, Sen, et al.
Published: (2025)
UniSync: A Unified Framework for Audio-Visual Synchronization
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning
by: Cheng, Xin, et al.
Published: (2026)
by: Cheng, Xin, et al.
Published: (2026)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
StereoSync: Spatially-Aware Stereo Audio Generation from Video
by: Marinoni, Christian, et al.
Published: (2025)
by: Marinoni, Christian, et al.
Published: (2025)
Style-Preserving Lip Sync via Audio-Aware Style Reference
by: Zhong, Weizhi, et al.
Published: (2024)
by: Zhong, Weizhi, et al.
Published: (2024)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes
by: Liu, Weifeng, et al.
Published: (2024)
by: Liu, Weifeng, et al.
Published: (2024)
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025)
by: Weng, Shuchen, et al.
Published: (2025)
OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
by: Zhang, Wenli, et al.
Published: (2026)
by: Zhang, Wenli, et al.
Published: (2026)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
by: Li, Chunyu, et al.
Published: (2024)
by: Li, Chunyu, et al.
Published: (2024)
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Fast-Slow Co-advancing Optimizer: Toward Harmonious Adversarial Training of GAN
by: Wang, Lin, et al.
Published: (2025)
by: Wang, Lin, et al.
Published: (2025)
Native Audio-Visual Alignment for Generation
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
Visual Harmony: Text-Visual Interplay in Circular Infographics
by: He, Shuqi, et al.
Published: (2024)
by: He, Shuqi, et al.
Published: (2024)
AnchorSync: Global Consistency Optimization for Long Video Editing
by: Liu, Zichi, et al.
Published: (2025)
by: Liu, Zichi, et al.
Published: (2025)
Q-Adapt: Adapting LMM for Visual Quality Assessment with Progressive Instruction Tuning
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
by: Li, Bingchen, et al.
Published: (2024)
by: Li, Bingchen, et al.
Published: (2024)
Hybrid Agents for Image Restoration
by: Li, Bingchen, et al.
Published: (2025)
by: Li, Bingchen, et al.
Published: (2025)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
by: Peng, Xinge, et al.
Published: (2026)
by: Peng, Xinge, et al.
Published: (2026)
Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances
by: Lu, Shilin, et al.
Published: (2024)
by: Lu, Shilin, et al.
Published: (2024)
Not in Sync: Unveiling Temporal Bias in Audio Chat Models
by: Yao, Jiayu, et al.
Published: (2025)
by: Yao, Jiayu, et al.
Published: (2025)
Reward-Aware Trajectory Shaping for Few-step Visual Generation
by: Li, Rui, et al.
Published: (2026)
by: Li, Rui, et al.
Published: (2026)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
by: Araujo, Edson, et al.
Published: (2025)
by: Araujo, Edson, et al.
Published: (2025)
ReTokSync: Self-Synchronizing Tokenization Disambiguation for Generative Linguistic Steganography
by: Wang, Yaofei, et al.
Published: (2026)
by: Wang, Yaofei, et al.
Published: (2026)
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
by: Li, Maomao, et al.
Published: (2026)
by: Li, Maomao, et al.
Published: (2026)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
by: Liu, Haohe, et al.
Published: (2024)
by: Liu, Haohe, et al.
Published: (2024)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
by: Ling, Zeyu, et al.
Published: (2025)
by: Ling, Zeyu, et al.
Published: (2025)
SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation
by: Wang, Hongrui, et al.
Published: (2026)
by: Wang, Hongrui, et al.
Published: (2026)
Object-AVEdit: An Object-level Audio-Visual Editing Model
by: Fu, Youquan, et al.
Published: (2025)
by: Fu, Youquan, et al.
Published: (2025)
Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
by: Liu, Lin, et al.
Published: (2026)
by: Liu, Lin, et al.
Published: (2026)
Similar Items
-
Video Quality Assessment Based on Swin TransformerV2 and Coarse to Fine Strategy
by: Yu, Zihao, et al.
Published: (2024) -
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model
by: Guan, Fengbin, et al.
Published: (2025) -
QMamba: On First Exploration of Vision Mamba for Image Quality Assessment
by: Guan, Fengbin, et al.
Published: (2024) -
Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
by: Hu, Teng, et al.
Published: (2025) -
OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
by: Liang, Sen, et al.
Published: (2025)