Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Xinxuan, Fowlkes, Charless, Berg, Alexander C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CriSp: Leveraging Tread Depth Maps for Enhanced Crime-Scene Shoeprint Matching
von: Shafique, Samia, et al.
Veröffentlicht: (2024)
von: Shafique, Samia, et al.
Veröffentlicht: (2024)
Instance Tracking in 3D Scenes from Egocentric Videos
von: Zhao, Yunhan, et al.
Veröffentlicht: (2023)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2023)
Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions
von: Kim, Kyungmin, et al.
Veröffentlicht: (2024)
von: Kim, Kyungmin, et al.
Veröffentlicht: (2024)
Polygon Intersection-over-Union Loss for Viewpoint-Agnostic Monocular 3D Vehicle Detection
von: Lu, Xinxuan, et al.
Veröffentlicht: (2023)
von: Lu, Xinxuan, et al.
Veröffentlicht: (2023)
Customizing Text-to-Image Diffusion with Object Viewpoint Control
von: Kumari, Nupur, et al.
Veröffentlicht: (2024)
von: Kumari, Nupur, et al.
Veröffentlicht: (2024)
ArbiViewGen: Controllable Arbitrary Viewpoint Camera Data Generation for Autonomous Driving via Stable Diffusion Models
von: Lan, Yatong, et al.
Veröffentlicht: (2025)
von: Lan, Yatong, et al.
Veröffentlicht: (2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
von: He, Hao, et al.
Veröffentlicht: (2024)
von: He, Hao, et al.
Veröffentlicht: (2024)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
von: Hou, Xinyu, et al.
Veröffentlicht: (2024)
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
von: Zhao, Zelin, et al.
Veröffentlicht: (2025)
von: Zhao, Zelin, et al.
Veröffentlicht: (2025)
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
von: Bai, Jianhong, et al.
Veröffentlicht: (2024)
von: Bai, Jianhong, et al.
Veröffentlicht: (2024)
Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic Field
von: Tang, Sheyang, et al.
Veröffentlicht: (2026)
von: Tang, Sheyang, et al.
Veröffentlicht: (2026)
Token Warping Helps MLLMs Look from Nearby Viewpoints
von: Lee, Phillip Y., et al.
Veröffentlicht: (2026)
von: Lee, Phillip Y., et al.
Veröffentlicht: (2026)
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets
von: Liu, Zhixuan, et al.
Veröffentlicht: (2026)
von: Liu, Zhixuan, et al.
Veröffentlicht: (2026)
PreciseCam: Precise Camera Control for Text-to-Image Generation
von: Bernal-Berdun, Edurne, et al.
Veröffentlicht: (2025)
von: Bernal-Berdun, Edurne, et al.
Veröffentlicht: (2025)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
von: Cao, Pu, et al.
Veröffentlicht: (2024)
von: Cao, Pu, et al.
Veröffentlicht: (2024)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
von: Wang, Zirui, et al.
Veröffentlicht: (2023)
von: Wang, Zirui, et al.
Veröffentlicht: (2023)
Aligning Text, Images, and 3D Structure Token-by-Token
von: Sahoo, Aadarsh, et al.
Veröffentlicht: (2025)
von: Sahoo, Aadarsh, et al.
Veröffentlicht: (2025)
RawGen: Learning Camera Raw Image Generation
von: Kim, Dongyoung, et al.
Veröffentlicht: (2026)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2026)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
Rethinking Attention-Based Multiple Instance Learning for Whole-Slide Pathological Image Classification: An Instance Attribute Viewpoint
von: Cai, Linghan, et al.
Veröffentlicht: (2024)
von: Cai, Linghan, et al.
Veröffentlicht: (2024)
Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimization
von: Li, Feifei, et al.
Veröffentlicht: (2025)
von: Li, Feifei, et al.
Veröffentlicht: (2025)
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
von: Wang, Yukun, et al.
Veröffentlicht: (2026)
von: Wang, Yukun, et al.
Veröffentlicht: (2026)
FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
von: Tang, Jiapeng, et al.
Veröffentlicht: (2025)
von: Tang, Jiapeng, et al.
Veröffentlicht: (2025)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
von: Han, Minghao, et al.
Veröffentlicht: (2025)
von: Han, Minghao, et al.
Veröffentlicht: (2025)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
High-Quality Virtual Single-Viewpoint Surgical Video: Geometric Autocalibration of Multiple Cameras in Surgical Lights
von: Kato, Yuna, et al.
Veröffentlicht: (2025)
von: Kato, Yuna, et al.
Veröffentlicht: (2025)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
von: Xie, Qingsong, et al.
Veröffentlicht: (2025)
von: Xie, Qingsong, et al.
Veröffentlicht: (2025)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
von: Yue, Yang, et al.
Veröffentlicht: (2026)
von: Yue, Yang, et al.
Veröffentlicht: (2026)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model
von: Zheng, Guangcong, et al.
Veröffentlicht: (2024)
von: Zheng, Guangcong, et al.
Veröffentlicht: (2024)
Powerful and Flexible: Personalized Text-to-Image Generation via Reinforcement Learning
von: Wei, Fanyue, et al.
Veröffentlicht: (2024)
von: Wei, Fanyue, et al.
Veröffentlicht: (2024)
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
von: Li, Teng, et al.
Veröffentlicht: (2025)
von: Li, Teng, et al.
Veröffentlicht: (2025)
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets
von: Coholich, Jeremiah, et al.
Veröffentlicht: (2026)
von: Coholich, Jeremiah, et al.
Veröffentlicht: (2026)
Local Representative Token Guided Merging for Text-to-Image Generation
von: Lee, Min-Jeong, et al.
Veröffentlicht: (2025)
von: Lee, Min-Jeong, et al.
Veröffentlicht: (2025)
Semantic-Aware Prefix Learning for Token-Efficient Image Generation
von: Li, Qingfeng, et al.
Veröffentlicht: (2026)
von: Li, Qingfeng, et al.
Veröffentlicht: (2026)
Training-free Camera Control for Video Generation
von: Hou, Chen, et al.
Veröffentlicht: (2024)
von: Hou, Chen, et al.
Veröffentlicht: (2024)
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
von: Jiang, Longtao, et al.
Veröffentlicht: (2025)
von: Jiang, Longtao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CriSp: Leveraging Tread Depth Maps for Enhanced Crime-Scene Shoeprint Matching
von: Shafique, Samia, et al.
Veröffentlicht: (2024) -
Instance Tracking in 3D Scenes from Egocentric Videos
von: Zhao, Yunhan, et al.
Veröffentlicht: (2023) -
Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions
von: Kim, Kyungmin, et al.
Veröffentlicht: (2024) -
Polygon Intersection-over-Union Loss for Viewpoint-Agnostic Monocular 3D Vehicle Detection
von: Lu, Xinxuan, et al.
Veröffentlicht: (2023) -
Customizing Text-to-Image Diffusion with Object Viewpoint Control
von: Kumari, Nupur, et al.
Veröffentlicht: (2024)