CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Haoyu, Zhang, Zihao, Gu, Jiaxi, Chen, Haoran, Zheng, Qingping, Tang, Pin, Jin, Yeyin, Zhang, Yuang, Cheng, Junqi, Lu, Zenghui, Shu, Peng, Wu, Zuxuan, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
ShoulderShot: Generating Over-the-Shoulder Dialogue Videos
von: Zhang, Yuang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuang, et al.
Veröffentlicht: (2025)
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023)
Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
von: Zhang, Xing, et al.
Veröffentlicht: (2024)
von: Zhang, Xing, et al.
Veröffentlicht: (2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
von: Zhao, Ruosen, et al.
Veröffentlicht: (2025)
von: Zhao, Ruosen, et al.
Veröffentlicht: (2025)
Predicting Camera Pose from Perspective Descriptions for Spatial Reasoning
von: Zhang, Xuejun, et al.
Veröffentlicht: (2026)
von: Zhang, Xuejun, et al.
Veröffentlicht: (2026)
Hybrid Spiking Vision Transformer for Object Detection with Event Cameras
von: Xu, Qi, et al.
Veröffentlicht: (2025)
von: Xu, Qi, et al.
Veröffentlicht: (2025)
Unify Robot Actions in Camera Frame
von: Xie, Sicheng, et al.
Veröffentlicht: (2025)
von: Xie, Sicheng, et al.
Veröffentlicht: (2025)
MotionMaster: Training-free Camera Motion Transfer For Video Generation
von: Hu, Teng, et al.
Veröffentlicht: (2024)
von: Hu, Teng, et al.
Veröffentlicht: (2024)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
Event USKT : U-State Space Model in Knowledge Transfer for Event Cameras
von: Lin, Yuhui, et al.
Veröffentlicht: (2024)
von: Lin, Yuhui, et al.
Veröffentlicht: (2024)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
von: He, Hao, et al.
Veröffentlicht: (2024)
von: He, Hao, et al.
Veröffentlicht: (2024)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2026)
von: Wu, Hang, et al.
Veröffentlicht: (2026)
Egocentric Gaze Estimation via Neck-Mounted Camera
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
von: Wu, Qiucheng, et al.
Veröffentlicht: (2025)
von: Wu, Qiucheng, et al.
Veröffentlicht: (2025)
Deep Unrolling Networks with Recurrent Momentum Acceleration for Nonlinear Inverse Problems
von: Zhou, Qingping, et al.
Veröffentlicht: (2023)
von: Zhou, Qingping, et al.
Veröffentlicht: (2023)
Enabling Cross-Camera Collaboration for Video Analytics on Distributed Smart Cameras
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
Paleoinspired Vision: From Exploring Colour Vision Evolution to Inspiring Camera Design
von: Zhang, Junjie, et al.
Veröffentlicht: (2024)
von: Zhang, Junjie, et al.
Veröffentlicht: (2024)
WHAC: World-grounded Humans and Cameras
von: Yin, Wanqi, et al.
Veröffentlicht: (2024)
von: Yin, Wanqi, et al.
Veröffentlicht: (2024)
High-speed and High-quality Vision Reconstruction of Spike Camera with Spike Stability Theorem
von: Zhang, Wei, et al.
Veröffentlicht: (2024)
von: Zhang, Wei, et al.
Veröffentlicht: (2024)
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)
von: Wei, Yana, et al.
Veröffentlicht: (2025)
CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
von: Wang, Yuelei, et al.
Veröffentlicht: (2024)
von: Wang, Yuelei, et al.
Veröffentlicht: (2024)
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
von: Lyu, Weijie, et al.
Veröffentlicht: (2026)
von: Lyu, Weijie, et al.
Veröffentlicht: (2026)
Adaptive Camera Sensor for Vision Models
von: Baek, Eunsu, et al.
Veröffentlicht: (2025)
von: Baek, Eunsu, et al.
Veröffentlicht: (2025)
Computer Vision with a Superpixelation Camera
von: Mahalingam, Sasidharan, et al.
Veröffentlicht: (2026)
von: Mahalingam, Sasidharan, et al.
Veröffentlicht: (2026)
Camera Obscura, Camera Lucida
Veröffentlicht: (2010)
Veröffentlicht: (2010)
I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
Unified Camera Positional Encoding for Controlled Video Generation
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
Probing into Camera Control of Video Models
von: Hou, Chen, et al.
Veröffentlicht: (2026)
von: Hou, Chen, et al.
Veröffentlicht: (2026)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
von: Wang, Cong, et al.
Veröffentlicht: (2024)
von: Wang, Cong, et al.
Veröffentlicht: (2024)
Seeing Through Pixel Motion: Learning Obstacle Avoidance from Optical Flow with One Camera
von: Hu, Yu, et al.
Veröffentlicht: (2024)
von: Hu, Yu, et al.
Veröffentlicht: (2024)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026) -
DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026) -
ShoulderShot: Generating Over-the-Shoulder Dialogue Videos
von: Zhang, Yuang, et al.
Veröffentlicht: (2025) -
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023) -
Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)