Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Zhening, Fox, Gereon, Zeng, Yanhong, Pan, Xingang, Elgharib, Mohamed, Theobalt, Christian, Chen, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic EventNeRF: Reconstructing General Dynamic Scenes from Multi-view RGB and Event Streams
by: Rudnev, Viktor, et al.
Published: (2024)
by: Rudnev, Viktor, et al.
Published: (2024)
EventNeuS: 3D Mesh Reconstruction from a Single Event Camera
by: Sachan, Shreyas, et al.
Published: (2026)
by: Sachan, Shreyas, et al.
Published: (2026)
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
Lite2Relight: 3D-aware Single Image Portrait Relighting
by: Rao, Pramod, et al.
Published: (2024)
by: Rao, Pramod, et al.
Published: (2024)
GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine Representations
by: Teotia, Kartik, et al.
Published: (2024)
by: Teotia, Kartik, et al.
Published: (2024)
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
DiffAge3D: Diffusion-based 3D-aware Face Aging
by: Wahid, Junaid, et al.
Published: (2024)
by: Wahid, Junaid, et al.
Published: (2024)
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
by: Chen, Joya, et al.
Published: (2025)
by: Chen, Joya, et al.
Published: (2025)
3DPR: Single Image 3D Portrait Relight using Generative Priors
by: Rao, Pramod, et al.
Published: (2025)
by: Rao, Pramod, et al.
Published: (2025)
LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
by: Yang, Zhenyu, et al.
Published: (2025)
by: Yang, Zhenyu, et al.
Published: (2025)
Online Misinformation Detection in Live Streaming Videos
by: Cao, Rui
Published: (2025)
by: Cao, Rui
Published: (2025)
Online Anomaly Detection over Live Social Video Streaming
by: He, Chengkun, et al.
Published: (2023)
by: He, Chengkun, et al.
Published: (2023)
PersonaLive! Expressive Portrait Image Animation for Live Streaming
by: Li, Zhiyuan, et al.
Published: (2025)
by: Li, Zhiyuan, et al.
Published: (2025)
Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
by: Zhao, Fangmin, et al.
Published: (2025)
by: Zhao, Fangmin, et al.
Published: (2025)
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
by: Ouyang, Wenqi, et al.
Published: (2024)
by: Ouyang, Wenqi, et al.
Published: (2024)
High-Quality Live Video Streaming via Transcoding Time Prediction and Preset Selection
by: Shahre-Babak, Zahra Nabizadeh, et al.
Published: (2023)
by: Shahre-Babak, Zahra Nabizadeh, et al.
Published: (2023)
LiveStre4m: Feed-Forward Live Streaming of Novel Views from Unposed Multi-View Video
by: Quesado, Pedro, et al.
Published: (2026)
by: Quesado, Pedro, et al.
Published: (2026)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
by: Ning, Zhenyu, et al.
Published: (2025)
by: Ning, Zhenyu, et al.
Published: (2025)
Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
by: Pan, Xingang, et al.
Published: (2023)
by: Pan, Xingang, et al.
Published: (2023)
Video Diffusion Models are Training-free Motion Interpreter and Controller
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
by: Shiu, Hau-Shiang, et al.
Published: (2025)
by: Shiu, Hau-Shiang, et al.
Published: (2025)
Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Live Video Captioning
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
by: Zou, Shilong, et al.
Published: (2025)
by: Zou, Shilong, et al.
Published: (2025)
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
by: Liu, Yanchen, et al.
Published: (2025)
by: Liu, Yanchen, et al.
Published: (2025)
LiveMoments: Reselected Key Photo Restoration in Live Photos via Reference-guided Diffusion
by: Xue, Clara, et al.
Published: (2026)
by: Xue, Clara, et al.
Published: (2026)
Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion
by: Chen, Zhifei, et al.
Published: (2024)
by: Chen, Zhifei, et al.
Published: (2024)
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
Trajectory Attention for Fine-grained Video Motion Control
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
by: Xie, Chaohao, et al.
Published: (2025)
by: Xie, Chaohao, et al.
Published: (2025)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
DiffSLT: Enhancing Diversity in Sign Language Translation via Diffusion Model
by: Moon, JiHwan, et al.
Published: (2024)
by: Moon, JiHwan, et al.
Published: (2024)
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
by: Zeng, Ziyun, et al.
Published: (2026)
by: Zeng, Ziyun, et al.
Published: (2026)
A Multimodal Transformer for Live Streaming Highlight Prediction
by: Deng, Jiaxin, et al.
Published: (2024)
by: Deng, Jiaxin, et al.
Published: (2024)
WorldMem: Long-term Consistent World Simulation with Memory
by: Xiao, Zeqi, et al.
Published: (2025)
by: Xiao, Zeqi, et al.
Published: (2025)
UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation
by: Liu, Zeyang, et al.
Published: (2025)
by: Liu, Zeyang, et al.
Published: (2025)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
by: Xia, Tian, et al.
Published: (2024)
by: Xia, Tian, et al.
Published: (2024)
Live Interactive Training for Video Segmentation
by: Yang, Xinyu, et al.
Published: (2026)
by: Yang, Xinyu, et al.
Published: (2026)
Similar Items
-
Dynamic EventNeRF: Reconstructing General Dynamic Scenes from Multi-view RGB and Event Streams
by: Rudnev, Viktor, et al.
Published: (2024) -
EventNeuS: 3D Mesh Reconstruction from a Single Event Camera
by: Sachan, Shreyas, et al.
Published: (2026) -
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023) -
Lite2Relight: 3D-aware Single Image Portrait Relighting
by: Rao, Pramod, et al.
Published: (2024) -
GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine Representations
by: Teotia, Kartik, et al.
Published: (2024)