Interactive Video Generation via Domain Adaptation
Fuente:
arXiv
Guardado en:
| Autores principales: | Rawal, Ishaan, Kumar, Suryansh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
por: Gu, Jing, et al.
Publicado: (2024)
por: Gu, Jing, et al.
Publicado: (2024)
Layer-wise Model Merging for Unsupervised Domain Adaptation in Segmentation Tasks
por: Alcover-Couso, Roberto, et al.
Publicado: (2024)
por: Alcover-Couso, Roberto, et al.
Publicado: (2024)
Image Conductor: Precision Control for Interactive Video Synthesis
por: Li, Yaowei, et al.
Publicado: (2024)
por: Li, Yaowei, et al.
Publicado: (2024)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
por: Huang, Feizhen, et al.
Publicado: (2025)
por: Huang, Feizhen, et al.
Publicado: (2025)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
por: Qing, Yuan, et al.
Publicado: (2026)
por: Qing, Yuan, et al.
Publicado: (2026)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
por: Zhou, Pengyuan, et al.
Publicado: (2024)
por: Zhou, Pengyuan, et al.
Publicado: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
por: Yan, Xin, et al.
Publicado: (2024)
por: Yan, Xin, et al.
Publicado: (2024)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024)
por: Zhang, Yiyuan, et al.
Publicado: (2024)
Advance Fake Video Detection via Vision Transformers
por: Battocchio, Joy, et al.
Publicado: (2025)
por: Battocchio, Joy, et al.
Publicado: (2025)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
por: Wu, Hao, et al.
Publicado: (2024)
por: Wu, Hao, et al.
Publicado: (2024)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
por: Cao, Pu, et al.
Publicado: (2023)
por: Cao, Pu, et al.
Publicado: (2023)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
por: Shin, Yosub, et al.
Publicado: (2025)
por: Shin, Yosub, et al.
Publicado: (2025)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
por: Qu, Qiang, et al.
Publicado: (2025)
por: Qu, Qiang, et al.
Publicado: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
por: Xiao, Xinyu, et al.
Publicado: (2026)
por: Xiao, Xinyu, et al.
Publicado: (2026)
Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery
por: Xu, Yulin, et al.
Publicado: (2026)
por: Xu, Yulin, et al.
Publicado: (2026)
Composing Concepts from Images and Videos via Concept-prompt Binding
por: Kong, Xianghao, et al.
Publicado: (2025)
por: Kong, Xianghao, et al.
Publicado: (2025)
AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos
por: Hu, Jiagao, et al.
Publicado: (2026)
por: Hu, Jiagao, et al.
Publicado: (2026)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
por: Zeng, Xiangyu, et al.
Publicado: (2024)
por: Zeng, Xiangyu, et al.
Publicado: (2024)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
por: Gan, Qijun, et al.
Publicado: (2025)
por: Gan, Qijun, et al.
Publicado: (2025)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
por: Wang, Yihao, et al.
Publicado: (2024)
por: Wang, Yihao, et al.
Publicado: (2024)
Regularizing Subspace Redundancy of Low-Rank Adaptation
por: Zhu, Yue, et al.
Publicado: (2025)
por: Zhu, Yue, et al.
Publicado: (2025)
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
por: Kim, Ye-Chan, et al.
Publicado: (2025)
por: Kim, Ye-Chan, et al.
Publicado: (2025)
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
por: Zhang, Yinghui, et al.
Publicado: (2025)
por: Zhang, Yinghui, et al.
Publicado: (2025)
Video Seal: Open and Efficient Video Watermarking
por: Fernandez, Pierre, et al.
Publicado: (2024)
por: Fernandez, Pierre, et al.
Publicado: (2024)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
por: Yuan, Hangjie, et al.
Publicado: (2025)
por: Yuan, Hangjie, et al.
Publicado: (2025)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
por: Zhang, Yuang, et al.
Publicado: (2024)
por: Zhang, Yuang, et al.
Publicado: (2024)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
por: Li, Ke, et al.
Publicado: (2026)
por: Li, Ke, et al.
Publicado: (2026)
EVAN: Evolutional Video Streaming Adaptation via Neural Representation
por: Liu, Mufan, et al.
Publicado: (2024)
por: Liu, Mufan, et al.
Publicado: (2024)
Programmable-Room: Interactive Textured 3D Room Meshes Generation Empowered by Large Language Models
por: Kim, Jihyun, et al.
Publicado: (2025)
por: Kim, Jihyun, et al.
Publicado: (2025)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
por: Tao, Ziyuan, et al.
Publicado: (2025)
por: Tao, Ziyuan, et al.
Publicado: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
por: Wang, Yun, et al.
Publicado: (2025)
por: Wang, Yun, et al.
Publicado: (2025)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
por: Meng, Jiahao, et al.
Publicado: (2025)
por: Meng, Jiahao, et al.
Publicado: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
por: Bian, Yuxuan, et al.
Publicado: (2025)
por: Bian, Yuxuan, et al.
Publicado: (2025)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
por: Liu, Jiajun, et al.
Publicado: (2024)
por: Liu, Jiajun, et al.
Publicado: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
por: Cai, Minghong, et al.
Publicado: (2024)
por: Cai, Minghong, et al.
Publicado: (2024)
Diffusion Models for Joint Audio-Video Generation
por: La Torre, Alejandro Paredes
Publicado: (2026)
por: La Torre, Alejandro Paredes
Publicado: (2026)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
por: Chen, Zhen, et al.
Publicado: (2025)
por: Chen, Zhen, et al.
Publicado: (2025)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
por: Ku, Max, et al.
Publicado: (2024)
por: Ku, Max, et al.
Publicado: (2024)
Question-Answering Dense Video Events
por: Qin, Hangyu, et al.
Publicado: (2024)
por: Qin, Hangyu, et al.
Publicado: (2024)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
por: Zhao, Ruixiang, et al.
Publicado: (2024)
por: Zhao, Ruixiang, et al.
Publicado: (2024)
Ejemplares similares
-
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
por: Gu, Jing, et al.
Publicado: (2024) -
Layer-wise Model Merging for Unsupervised Domain Adaptation in Segmentation Tasks
por: Alcover-Couso, Roberto, et al.
Publicado: (2024) -
Image Conductor: Precision Control for Interactive Video Synthesis
por: Li, Yaowei, et al.
Publicado: (2024) -
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
por: Huang, Feizhen, et al.
Publicado: (2025) -
Moiré Video Authentication: A Physical Signature Against AI Video Generation
por: Qing, Yuan, et al.
Publicado: (2026)