Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Subin, Oh, Seoung Wug, Wang, Jui-Hsien, Lee, Joon-Young, Shin, Jinwoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Elevating Flow-Guided Video Inpainting with Reference Generation
di: Cho, Suhwan, et al.
Pubblicazione: (2024)
di: Cho, Suhwan, et al.
Pubblicazione: (2024)
VideoMaMa: Mask-Guided Video Matting via Generative Prior
di: Lim, Sangbeom, et al.
Pubblicazione: (2026)
di: Lim, Sangbeom, et al.
Pubblicazione: (2026)
Putting the Object Back into Video Object Segmentation
di: Cheng, Ho Kei, et al.
Pubblicazione: (2023)
di: Cheng, Ho Kei, et al.
Pubblicazione: (2023)
MaGGIe: Masked Guided Gradual Human Instance Matting
di: Huynh, Chuong, et al.
Pubblicazione: (2024)
di: Huynh, Chuong, et al.
Pubblicazione: (2024)
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
di: TV, Sethuraman, et al.
Pubblicazione: (2025)
di: TV, Sethuraman, et al.
Pubblicazione: (2025)
IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation
di: Yang, Sejong, et al.
Pubblicazione: (2024)
di: Yang, Sejong, et al.
Pubblicazione: (2024)
HARIVO: Harnessing Text-to-Image Models for Video Generation
di: Kwon, Mingi, et al.
Pubblicazione: (2024)
di: Kwon, Mingi, et al.
Pubblicazione: (2024)
VISAGE: Video Instance Segmentation with Appearance-Guided Enhancement
di: Kim, Hanjung, et al.
Pubblicazione: (2023)
di: Kim, Hanjung, et al.
Pubblicazione: (2023)
DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
di: Ngo, Tuan Duc, et al.
Pubblicazione: (2026)
di: Ngo, Tuan Duc, et al.
Pubblicazione: (2026)
In-N-Out: Faithful 3D GAN Inversion with Volumetric Decomposition for Face Editing
di: Xu, Yiran, et al.
Pubblicazione: (2023)
di: Xu, Yiran, et al.
Pubblicazione: (2023)
Video Color Grading via Look-Up Table Generation
di: Shin, Seunghyun, et al.
Pubblicazione: (2025)
di: Shin, Seunghyun, et al.
Pubblicazione: (2025)
FontAdapter: Instant Font Adaptation in Visual Text Generation
di: Koo, Myungkyu, et al.
Pubblicazione: (2025)
di: Koo, Myungkyu, et al.
Pubblicazione: (2025)
Generative Video Motion Editing with 3D Point Tracks
di: Lee, Yao-Chih, et al.
Pubblicazione: (2025)
di: Lee, Yao-Chih, et al.
Pubblicazione: (2025)
Generative Video Propagation
di: Liu, Shaoteng, et al.
Pubblicazione: (2024)
di: Liu, Shaoteng, et al.
Pubblicazione: (2024)
ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative Adaptation
di: Youwang, Kim, et al.
Pubblicazione: (2026)
di: Youwang, Kim, et al.
Pubblicazione: (2026)
Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-Reasoning
di: Park, Subin, et al.
Pubblicazione: (2026)
di: Park, Subin, et al.
Pubblicazione: (2026)
Restoration-Aligned Generative Flow Models for Blind Motion Deblurring
di: Kim, Insoo, et al.
Pubblicazione: (2026)
di: Kim, Insoo, et al.
Pubblicazione: (2026)
Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
di: Lee, Kyungmin, et al.
Pubblicazione: (2025)
di: Lee, Kyungmin, et al.
Pubblicazione: (2025)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
di: Woo, Young Beom, et al.
Pubblicazione: (2025)
di: Woo, Young Beom, et al.
Pubblicazione: (2025)
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
di: Ma, Yongjia, et al.
Pubblicazione: (2025)
di: Ma, Yongjia, et al.
Pubblicazione: (2025)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
di: Oh, Ju-Young
Pubblicazione: (2025)
di: Oh, Ju-Young
Pubblicazione: (2025)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
di: Bond, Andrew, et al.
Pubblicazione: (2025)
di: Bond, Andrew, et al.
Pubblicazione: (2025)
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
di: Lu, Yu, et al.
Pubblicazione: (2025)
di: Lu, Yu, et al.
Pubblicazione: (2025)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
di: Jeon, Subin, et al.
Pubblicazione: (2024)
di: Jeon, Subin, et al.
Pubblicazione: (2024)
PDF-GS: Progressive Distractor Filtering for Robust 3D Gaussian Splatting
di: Seo, Kangmin, et al.
Pubblicazione: (2026)
di: Seo, Kangmin, et al.
Pubblicazione: (2026)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
di: Hyun, Jeongseok, et al.
Pubblicazione: (2025)
di: Hyun, Jeongseok, et al.
Pubblicazione: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
di: Jang, Huiwon, et al.
Pubblicazione: (2024)
di: Jang, Huiwon, et al.
Pubblicazione: (2024)
CoAPT: Context Attribute words for Prompt Tuning
di: Lee, Gun, et al.
Pubblicazione: (2024)
di: Lee, Gun, et al.
Pubblicazione: (2024)
DreamFlow: High-Quality Text-to-3D Generation by Approximating Probability Flow
di: Lee, Kyungmin, et al.
Pubblicazione: (2024)
di: Lee, Kyungmin, et al.
Pubblicazione: (2024)
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets
di: Liu, Zhixuan, et al.
Pubblicazione: (2026)
di: Liu, Zhixuan, et al.
Pubblicazione: (2026)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
di: Choi, June Suk, et al.
Pubblicazione: (2025)
di: Choi, June Suk, et al.
Pubblicazione: (2025)
MEVG: Multi-event Video Generation with Text-to-Video Models
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding
di: Kim, Sunoh, et al.
Pubblicazione: (2023)
di: Kim, Sunoh, et al.
Pubblicazione: (2023)
HarmoVid: Relightful Video Portrait Harmonization
di: Choi, Jun Myeong, et al.
Pubblicazione: (2026)
di: Choi, Jun Myeong, et al.
Pubblicazione: (2026)
Long Context Tuning for Video Generation
di: Guo, Yuwei, et al.
Pubblicazione: (2025)
di: Guo, Yuwei, et al.
Pubblicazione: (2025)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
di: Park, Seong Hyeon, et al.
Pubblicazione: (2025)
di: Park, Seong Hyeon, et al.
Pubblicazione: (2025)
Iterative Prompt Refinement for Safer Text-to-Image Generation
di: Jeon, Jinwoo, et al.
Pubblicazione: (2025)
di: Jeon, Jinwoo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Elevating Flow-Guided Video Inpainting with Reference Generation
di: Cho, Suhwan, et al.
Pubblicazione: (2024) -
VideoMaMa: Mask-Guided Video Matting via Generative Prior
di: Lim, Sangbeom, et al.
Pubblicazione: (2026) -
Putting the Object Back into Video Object Segmentation
di: Cheng, Ho Kei, et al.
Pubblicazione: (2023) -
MaGGIe: Masked Guided Gradual Human Instance Matting
di: Huynh, Chuong, et al.
Pubblicazione: (2024) -
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
di: TV, Sethuraman, et al.
Pubblicazione: (2025)