CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
Fuente:
arXiv
Saved in:
| Main Authors: | Zi, Bojia, Zhao, Shihao, Qi, Xianbiao, Wang, Jianan, Shi, Yukai, Chen, Qianyu, Liang, Bin, Wong, Kam-Fai, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
by: Zi, Bojia, et al.
Published: (2025)
by: Zi, Bojia, et al.
Published: (2025)
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
by: Zi, Bojia, et al.
Published: (2025)
by: Zi, Bojia, et al.
Published: (2025)
DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation
by: Huang, Yukun, et al.
Published: (2023)
by: Huang, Yukun, et al.
Published: (2023)
Refaçade: Editing Object with Given Reference Texture
by: Huang, Youze, et al.
Published: (2025)
by: Huang, Youze, et al.
Published: (2025)
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
by: Hao, Shaozhe, et al.
Published: (2024)
by: Hao, Shaozhe, et al.
Published: (2024)
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
by: Lin, Jiahao, et al.
Published: (2025)
by: Lin, Jiahao, et al.
Published: (2025)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
CoCoDiff: Diversifying Skeleton Action Features via Coarse-Fine Text-Co-Guided Latent Diffusion
by: Zhao, Zhifu, et al.
Published: (2025)
by: Zhao, Zhifu, et al.
Published: (2025)
Ctrl&Shift: High-Quality Geometry-Aware Object Manipulation in Visual Generation
by: Ruan, Penghui, et al.
Published: (2026)
by: Ruan, Penghui, et al.
Published: (2026)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
by: Wang, Zezhong, et al.
Published: (2024)
by: Wang, Zezhong, et al.
Published: (2024)
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
by: Gao, Zhanxin, et al.
Published: (2025)
by: Gao, Zhanxin, et al.
Published: (2025)
3D Gaussian Inpainting with Depth-Guided Cross-View Consistency
by: Huang, Sheng-Yu, et al.
Published: (2025)
by: Huang, Sheng-Yu, et al.
Published: (2025)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising
by: Lyu, Shuangquan, et al.
Published: (2025)
by: Lyu, Shuangquan, et al.
Published: (2025)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
by: Han, Gyuwon, et al.
Published: (2026)
by: Han, Gyuwon, et al.
Published: (2026)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
by: Gong, Meiqi, et al.
Published: (2025)
by: Gong, Meiqi, et al.
Published: (2025)
CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer
by: Nie, Wenbo, et al.
Published: (2026)
by: Nie, Wenbo, et al.
Published: (2026)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
by: Mao, Zhiming, et al.
Published: (2024)
by: Mao, Zhiming, et al.
Published: (2024)
SparrowSNN: A Hardware/software Co-design for Energy Efficient ECG Classification
by: Yan, Zhanglu, et al.
Published: (2024)
by: Yan, Zhanglu, et al.
Published: (2024)
LLM-Guided Co-Training for Text Classification
by: Rahman, Md Mezbaur, et al.
Published: (2025)
by: Rahman, Md Mezbaur, et al.
Published: (2025)
WHERE and WHICH: Iterative Debate for Biomedical Synthetic Data Augmentation
by: Zhao, Zhengyi, et al.
Published: (2025)
by: Zhao, Zhengyi, et al.
Published: (2025)
Taming Transformer Without Using Learning Rate Warmup
by: Qi, Xianbiao, et al.
Published: (2025)
by: Qi, Xianbiao, et al.
Published: (2025)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning
by: Wu, Yuhui, et al.
Published: (2026)
by: Wu, Yuhui, et al.
Published: (2026)
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
by: Luo, Xiangyang, et al.
Published: (2026)
by: Luo, Xiangyang, et al.
Published: (2026)
CoTracker: It is Better to Track Together
by: Karaev, Nikita, et al.
Published: (2023)
by: Karaev, Nikita, et al.
Published: (2023)
PEARL: Towards Permutation-Resilient LLMs
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Semantically Consistent Video Inpainting with Conditional Diffusion Models
by: Green, Dylan, et al.
Published: (2024)
by: Green, Dylan, et al.
Published: (2024)
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
by: Li, Haodong, et al.
Published: (2026)
by: Li, Haodong, et al.
Published: (2026)
Nonlinear optical analogues of quantum phase transitions in a squeezing-enhanced LMG model
by: Kam, Chon-Fai
Published: (2025)
by: Kam, Chon-Fai
Published: (2025)
Majorana Constellations: A Geometric Lens on Multipartite Entanglement and Geometric Phases
by: Kam, Chon-Fai
Published: (2026)
by: Kam, Chon-Fai
Published: (2026)
Three-Axis Spin Squeezed States Associated with Excited-State Quantum Phase Transitions
by: Kam, Chon-Fai
Published: (2025)
by: Kam, Chon-Fai
Published: (2025)
Nonlinear optical realization of non-integrable phases accompanying quantum phase transitions
by: Kam, Chon-Fai
Published: (2025)
by: Kam, Chon-Fai
Published: (2025)
Two Teachers Better Than One: Hardware-Physics Co-Guided Distributed Scientific Machine Learning
by: Yuan, Yuchen, et al.
Published: (2026)
by: Yuan, Yuchen, et al.
Published: (2026)
CoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent Cooperation
by: Li, Renhao, et al.
Published: (2024)
by: Li, Renhao, et al.
Published: (2024)
CoNo: Consistency Noise Injection for Tuning-free Long Video Diffusion
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
by: Zeng, Qinglin, et al.
Published: (2025)
by: Zeng, Qinglin, et al.
Published: (2025)
CoCo-CoLa: Evaluating and Improving Language Adherence in Multilingual LLMs
by: Rahmati, Elnaz, et al.
Published: (2025)
by: Rahmati, Elnaz, et al.
Published: (2025)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
by: Qi, Ji, et al.
Published: (2023)
by: Qi, Ji, et al.
Published: (2023)
Bridging the Macro, Meso and Micro Levels of Designer–Artisan Co‐Design: Case Studies of Value Co‐Creation Within Chinese Textile Crafts
by: Jianan Hu, et al.
Published: (2025)
by: Jianan Hu, et al.
Published: (2025)
Similar Items
-
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
by: Zi, Bojia, et al.
Published: (2025) -
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
by: Zi, Bojia, et al.
Published: (2025) -
DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation
by: Huang, Yukun, et al.
Published: (2023) -
Refaçade: Editing Object with Given Reference Texture
by: Huang, Youze, et al.
Published: (2025) -
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
by: Hao, Shaozhe, et al.
Published: (2024)