Layer-Aware Video Composition via Split-then-Merge
Fuente:
arXiv
Salvato in:
| Autori principali: | Kara, Ozgur, Chen, Yujia, Yang, Ming-Hsuan, Rehg, James M., Chu, Wen-Sheng, Tran, Du |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images
di: Kara, Ozgur, et al.
Pubblicazione: (2025)
di: Kara, Ozgur, et al.
Pubblicazione: (2025)
Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation
di: Long, Zeqian, et al.
Pubblicazione: (2026)
di: Long, Zeqian, et al.
Pubblicazione: (2026)
SEAL: Semantic Attention Learning for Long Video Representation
di: Wang, Lan, et al.
Pubblicazione: (2024)
di: Wang, Lan, et al.
Pubblicazione: (2024)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
di: Kara, Ozgur, et al.
Pubblicazione: (2025)
di: Kara, Ozgur, et al.
Pubblicazione: (2025)
FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
di: Lyu, Weijie, et al.
Pubblicazione: (2026)
di: Lyu, Weijie, et al.
Pubblicazione: (2026)
DiffVax: Optimization-Free Image Immunization Against Diffusion-Based Editing
di: Ozden, Tarik Can, et al.
Pubblicazione: (2024)
di: Ozden, Tarik Can, et al.
Pubblicazione: (2024)
ToSA: Token Merging with Spatial Awareness
di: Huang, Hsiang-Wei, et al.
Pubblicazione: (2025)
di: Huang, Hsiang-Wei, et al.
Pubblicazione: (2025)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
di: Chu, Wen-Hsuan, et al.
Pubblicazione: (2024)
di: Chu, Wen-Hsuan, et al.
Pubblicazione: (2024)
Streaming Autoregressive Video Generation via Diagonal Distillation
di: Liu, Jinxiu, et al.
Pubblicazione: (2026)
di: Liu, Jinxiu, et al.
Pubblicazione: (2026)
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
di: Li, Xinyao, et al.
Pubblicazione: (2024)
di: Li, Xinyao, et al.
Pubblicazione: (2024)
Leveraging Object Priors for Point Tracking
di: Boote, Bikram, et al.
Pubblicazione: (2024)
di: Boote, Bikram, et al.
Pubblicazione: (2024)
Unified Dense Prediction of Video Diffusion
di: Yang, Lehan, et al.
Pubblicazione: (2025)
di: Yang, Lehan, et al.
Pubblicazione: (2025)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
di: Jiang, Xixi, et al.
Pubblicazione: (2025)
di: Jiang, Xixi, et al.
Pubblicazione: (2025)
How Much 3D Do Video Foundation Models Encode?
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
di: Wu, Hang, et al.
Pubblicazione: (2026)
di: Wu, Hang, et al.
Pubblicazione: (2026)
From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition
di: Lo, Ling, et al.
Pubblicazione: (2025)
di: Lo, Ling, et al.
Pubblicazione: (2025)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
di: Lai, Bolin, et al.
Pubblicazione: (2023)
di: Lai, Bolin, et al.
Pubblicazione: (2023)
Enhancing Quantization-Aware Training on Edge Devices via Relative Entropy Coreset Selection and Cascaded Layer Correction
di: Tong, Yujia, et al.
Pubblicazione: (2025)
di: Tong, Yujia, et al.
Pubblicazione: (2025)
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
di: Bao, Jingzhi, et al.
Pubblicazione: (2024)
di: Bao, Jingzhi, et al.
Pubblicazione: (2024)
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
di: Quattrini, Fabio, et al.
Pubblicazione: (2024)
di: Quattrini, Fabio, et al.
Pubblicazione: (2024)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
di: Huang, Kuan-Chih, et al.
Pubblicazione: (2023)
di: Huang, Kuan-Chih, et al.
Pubblicazione: (2023)
VideoMerge: Towards Training-free Long Video Generation
di: Zhang, Siyang, et al.
Pubblicazione: (2025)
di: Zhang, Siyang, et al.
Pubblicazione: (2025)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
di: Lai, Bolin, et al.
Pubblicazione: (2025)
di: Lai, Bolin, et al.
Pubblicazione: (2025)
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
di: Liu, Jinxiu, et al.
Pubblicazione: (2024)
di: Liu, Jinxiu, et al.
Pubblicazione: (2024)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
di: Kim, Junho, et al.
Pubblicazione: (2026)
di: Kim, Junho, et al.
Pubblicazione: (2026)
MambaVideo for Discrete Video Tokenization with Channel-Split Quantization
di: Argaw, Dawit Mureja, et al.
Pubblicazione: (2025)
di: Argaw, Dawit Mureja, et al.
Pubblicazione: (2025)
Vinedresser3D: Agentic Text-guided 3D Editing
di: Chi, Yankuan, et al.
Pubblicazione: (2026)
di: Chi, Yankuan, et al.
Pubblicazione: (2026)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
di: Lai, Bolin, et al.
Pubblicazione: (2022)
di: Lai, Bolin, et al.
Pubblicazione: (2022)
Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Towards Affordance-Aware Articulation Synthesis for Rigged Objects
di: Yu, Yu-Chu, et al.
Pubblicazione: (2025)
di: Yu, Yu-Chu, et al.
Pubblicazione: (2025)
One-for-All: Towards Universal Domain Translation with a Single StyleGAN
di: Du, Yong, et al.
Pubblicazione: (2023)
di: Du, Yong, et al.
Pubblicazione: (2023)
Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning
di: Jiang, Shengqin, et al.
Pubblicazione: (2025)
di: Jiang, Shengqin, et al.
Pubblicazione: (2025)
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
di: Fu, Zihang, et al.
Pubblicazione: (2026)
di: Fu, Zihang, et al.
Pubblicazione: (2026)
VSViG: Real-time Video-based Seizure Detection via Skeleton-based Spatiotemporal ViG
di: Xu, Yankun, et al.
Pubblicazione: (2023)
di: Xu, Yankun, et al.
Pubblicazione: (2023)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images
di: Kara, Ozgur, et al.
Pubblicazione: (2025) -
Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation
di: Long, Zeqian, et al.
Pubblicazione: (2026) -
SEAL: Semantic Attention Learning for Long Video Representation
di: Wang, Lan, et al.
Pubblicazione: (2024) -
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
di: Kara, Ozgur, et al.
Pubblicazione: (2025) -
FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
di: Lyu, Weijie, et al.
Pubblicazione: (2026)