OmniCam: Unified Multimodal Video Generation via Camera Control
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Xiaoda, Xu, Jiayang, Luan, Kaixuan, Zhan, Xinyu, Qiu, Hongshun, Shi, Shijun, Li, Hao, Yang, Shuai, Zhang, Li, Yu, Checheng, Lu, Cewu, Yang, Lixin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
by: Wang, Yukun, et al.
Published: (2026)
by: Wang, Yukun, et al.
Published: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control
by: Kuang, Zhiyi, et al.
Published: (2026)
by: Kuang, Zhiyi, et al.
Published: (2026)
ReCamDriving: LiDAR-Free Camera-Controlled Novel Trajectory Video Generation
by: Li, Yaokun, et al.
Published: (2025)
by: Li, Yaokun, et al.
Published: (2025)
FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
by: Lyu, Weijie, et al.
Published: (2026)
by: Lyu, Weijie, et al.
Published: (2026)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Motion Before Action: Diffusing Object Motion as Manipulation Condition
by: Su, Yue, et al.
Published: (2024)
by: Su, Yue, et al.
Published: (2024)
A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning
by: Yang, Xiaoda, et al.
Published: (2026)
by: Yang, Xiaoda, et al.
Published: (2026)
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
by: Zhang, Hongfei, et al.
Published: (2025)
by: Zhang, Hongfei, et al.
Published: (2025)
VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
by: Wu, Qiucheng, et al.
Published: (2025)
by: Wu, Qiucheng, et al.
Published: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
by: Liu, Jialun, et al.
Published: (2026)
by: Liu, Jialun, et al.
Published: (2026)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model
by: Zheng, Guangcong, et al.
Published: (2024)
by: Zheng, Guangcong, et al.
Published: (2024)
MemCam: Memory-Augmented Camera Control for Consistent Video Generation
by: Gao, Xinhang, et al.
Published: (2026)
by: Gao, Xinhang, et al.
Published: (2026)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
by: Ge, Wenhang, et al.
Published: (2026)
by: Ge, Wenhang, et al.
Published: (2026)
CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
by: Yang, Qirui, et al.
Published: (2025)
by: Yang, Qirui, et al.
Published: (2025)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
CamCloneMaster: Enabling Reference-based Camera Control for Video Generation
by: Luo, Yawen, et al.
Published: (2025)
by: Luo, Yawen, et al.
Published: (2025)
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
by: Li, Teng, et al.
Published: (2025)
by: Li, Teng, et al.
Published: (2025)
PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention
by: Chen, Yipeng, et al.
Published: (2025)
by: Chen, Yipeng, et al.
Published: (2025)
Multi-view Hand Reconstruction with a Point-Embedded Transformer
by: Yang, Lixin, et al.
Published: (2024)
by: Yang, Lixin, et al.
Published: (2024)
Dense Policy: Bidirectional Autoregressive Learning of Actions
by: Su, Yue, et al.
Published: (2025)
by: Su, Yue, et al.
Published: (2025)
Topology Optimization Design of Automotive Suspension Control Arm
by: Rongfeng Lin, et al.
Published: (2025)
by: Rongfeng Lin, et al.
Published: (2025)
OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion
by: Zhan, Xinyu, et al.
Published: (2024)
by: Zhan, Xinyu, et al.
Published: (2024)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
by: Marmon, Andrew, et al.
Published: (2024)
by: Marmon, Andrew, et al.
Published: (2024)
RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control
by: Xu, Youcan, et al.
Published: (2026)
by: Xu, Youcan, et al.
Published: (2026)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
by: He, Hao, et al.
Published: (2024)
by: He, Hao, et al.
Published: (2024)
Tri-Cam: Practical Eye Gaze Tracking via Camera Network
by: Yang, Sikai, et al.
Published: (2024)
by: Yang, Sikai, et al.
Published: (2024)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
by: Ji, Chenhao, et al.
Published: (2025)
by: Ji, Chenhao, et al.
Published: (2025)
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
by: Jiang, Hong, et al.
Published: (2026)
by: Jiang, Hong, et al.
Published: (2026)
SemGrasp: Semantic Grasp Generation via Language Aligned Discretization
by: Li, Kailin, et al.
Published: (2024)
by: Li, Kailin, et al.
Published: (2024)
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction
by: Xiong, Kai, et al.
Published: (2026)
by: Xiong, Kai, et al.
Published: (2026)
DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control
by: Xu, Xinyu, et al.
Published: (2024)
by: Xu, Xinyu, et al.
Published: (2024)
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
by: Liu, Haofeng, et al.
Published: (2026)
by: Liu, Haofeng, et al.
Published: (2026)
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
by: Bai, Jianhong, et al.
Published: (2025)
by: Bai, Jianhong, et al.
Published: (2025)
Astrea: A MOE-based Visual Understanding Model with Progressive Alignment
by: Yang, Xiaoda, et al.
Published: (2025)
by: Yang, Xiaoda, et al.
Published: (2025)
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
by: Nam, Jisu, et al.
Published: (2026)
by: Nam, Jisu, et al.
Published: (2026)
Similar Items
-
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
by: Wang, Yukun, et al.
Published: (2026) -
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025) -
CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control
by: Kuang, Zhiyi, et al.
Published: (2026) -
ReCamDriving: LiDAR-Free Camera-Controlled Novel Trajectory Video Generation
by: Li, Yaokun, et al.
Published: (2025) -
FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
by: Lyu, Weijie, et al.
Published: (2026)