PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Heyuan, Tang, Bangxun, Song, Yiren, Fang, Guian, He, Zijian, Yang, Jie, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
Mitty: Diffusion-based Human-to-Robot Video Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
von: Gu, Yuchao, et al.
Veröffentlicht: (2026)
von: Gu, Yuchao, et al.
Veröffentlicht: (2026)
Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
von: Shi, Yiqing, et al.
Veröffentlicht: (2025)
von: Shi, Yiqing, et al.
Veröffentlicht: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
von: Song, Yiren, et al.
Veröffentlicht: (2024)
von: Song, Yiren, et al.
Veröffentlicht: (2024)
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
von: Yang, Pei, et al.
Veröffentlicht: (2026)
von: Yang, Pei, et al.
Veröffentlicht: (2026)
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
von: Ci, Hai, et al.
Veröffentlicht: (2025)
von: Ci, Hai, et al.
Veröffentlicht: (2025)
Steganalysis on Digital Watermarking: Is Your Defense Truly Impervious?
von: Yang, Pei, et al.
Veröffentlicht: (2024)
von: Yang, Pei, et al.
Veröffentlicht: (2024)
IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
von: Song, Yiren, et al.
Veröffentlicht: (2024)
von: Song, Yiren, et al.
Veröffentlicht: (2024)
RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification
von: Ci, Hai, et al.
Veröffentlicht: (2024)
von: Ci, Hai, et al.
Veröffentlicht: (2024)
FramePrompt: In-context Controllable Animation with Zero Structural Changes
von: Fang, Guian, et al.
Veröffentlicht: (2025)
von: Fang, Guian, et al.
Veröffentlicht: (2025)
OmniPSD: Layered PSD Generation with Diffusion Transformer
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
ANYPORTAL: Zero-Shot Consistent Video Background Replacement
von: Gao, Wenshuo, et al.
Veröffentlicht: (2025)
von: Gao, Wenshuo, et al.
Veröffentlicht: (2025)
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
von: Ci, Hai, et al.
Veröffentlicht: (2024)
von: Ci, Hai, et al.
Veröffentlicht: (2024)
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
Personalized Vision via Visual In-Context Learning
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
von: Ran, Lingmin, et al.
Veröffentlicht: (2025)
von: Ran, Lingmin, et al.
Veröffentlicht: (2025)
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
FlowPortal: Residual-Corrected Flow for Training-Free Video Relighting and Background Replacement
von: Gao, Wenshuo, et al.
Veröffentlicht: (2025)
von: Gao, Wenshuo, et al.
Veröffentlicht: (2025)
Motion-Aware Optical Camera Communication with Event Cameras
von: Su, Hang, et al.
Veröffentlicht: (2024)
von: Su, Hang, et al.
Veröffentlicht: (2024)
MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation
von: Xing, Jinbo, et al.
Veröffentlicht: (2025)
von: Xing, Jinbo, et al.
Veröffentlicht: (2025)
Impossible Videos
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
von: Zeng, Jianshu, et al.
Veröffentlicht: (2025)
von: Zeng, Jianshu, et al.
Veröffentlicht: (2025)
D-AR: Diffusion via Autoregressive Models
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation
von: Song, Yiren, et al.
Veröffentlicht: (2024)
von: Song, Yiren, et al.
Veröffentlicht: (2024)
AssistPDA: An Online Video Surveillance Assistant for Video Anomaly Prediction, Detection, and Analysis
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
von: Gu, Yuchao, et al.
Veröffentlicht: (2025)
von: Gu, Yuchao, et al.
Veröffentlicht: (2025)
SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control
von: Zhang, Zhida, et al.
Veröffentlicht: (2026)
von: Zhang, Zhida, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
von: Song, Yiren, et al.
Veröffentlicht: (2026) -
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
von: Yang, Pei, et al.
Veröffentlicht: (2025) -
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026) -
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
von: Song, Quanjian, et al.
Veröffentlicht: (2025) -
Mitty: Diffusion-based Human-to-Robot Video Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)