MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Buyu, Wang, Kai, Liu, Yansong, Bao, Jun, Han, Tingting, Yu, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
von: Wang, Bing, et al.
Veröffentlicht: (2025)
von: Wang, Bing, et al.
Veröffentlicht: (2025)
Deep Contrastive Multi-view Clustering under Semantic Feature Guidance
von: Liu, Siwen, et al.
Veröffentlicht: (2024)
von: Liu, Siwen, et al.
Veröffentlicht: (2024)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
von: Liu, Ke, et al.
Veröffentlicht: (2025)
von: Liu, Ke, et al.
Veröffentlicht: (2025)
RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
Simple Yet Effective Selective Imputation for Incomplete Multi-view Clustering
von: Xu, Cai, et al.
Veröffentlicht: (2025)
von: Xu, Cai, et al.
Veröffentlicht: (2025)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
Accelerating Controllable Generation via Hybrid-grained Cache
von: Liu, Lin, et al.
Veröffentlicht: (2025)
von: Liu, Lin, et al.
Veröffentlicht: (2025)
Generalizable Deepfake Detection Based on Forgery-aware Layer Masking and Multi-artifact Subspace Decomposition
von: Zhang, Xiang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiang, et al.
Veröffentlicht: (2026)
Test-time adaptation for image compression with distribution regularization
von: Chen, Kecheng, et al.
Veröffentlicht: (2024)
von: Chen, Kecheng, et al.
Veröffentlicht: (2024)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
von: Wu, Yi, et al.
Veröffentlicht: (2025)
von: Wu, Yi, et al.
Veröffentlicht: (2025)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
von: Zhang, Wang, et al.
Veröffentlicht: (2024)
von: Zhang, Wang, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
T$^\text{3}$SVFND: Towards an Evolving Fake News Detector for Emergencies with Test-time Training on Short Video Platforms
von: Zhang, Liyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Liyuan, et al.
Veröffentlicht: (2025)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis
von: Liu, Miao, et al.
Veröffentlicht: (2026)
von: Liu, Miao, et al.
Veröffentlicht: (2026)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
G-Refine: A General Quality Refiner for Text-to-Image Generation
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
von: Yuan, Bo, et al.
Veröffentlicht: (2024)
von: Yuan, Bo, et al.
Veröffentlicht: (2024)
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
von: Yang, Rui, et al.
Veröffentlicht: (2024)
von: Yang, Rui, et al.
Veröffentlicht: (2024)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2026)
von: Cai, Qi, et al.
Veröffentlicht: (2026)
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022)
von: Tang, Anni, et al.
Veröffentlicht: (2022)
Regularized Contrastive Partial Multi-view Outlier Detection
von: Wang, Yijia, et al.
Veröffentlicht: (2024)
von: Wang, Yijia, et al.
Veröffentlicht: (2024)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
Querying Autonomous Vehicle Point Clouds: Enhanced by 3D Object Counting with CounterNet
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
On-the-Fly Object-aware Representative Point Selection in Point Cloud
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
von: Wang, Xue, et al.
Veröffentlicht: (2026)
von: Wang, Xue, et al.
Veröffentlicht: (2026)
A Perspective on Deep Vision Performance with Standard Image and Video Codecs
von: Reich, Christoph, et al.
Veröffentlicht: (2024)
von: Reich, Christoph, et al.
Veröffentlicht: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
von: Wang, Bing, et al.
Veröffentlicht: (2025) -
Deep Contrastive Multi-view Clustering under Semantic Feature Guidance
von: Liu, Siwen, et al.
Veröffentlicht: (2024) -
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
von: Liu, Ke, et al.
Veröffentlicht: (2025) -
RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction
von: Liu, Shuhong, et al.
Veröffentlicht: (2025) -
Simple Yet Effective Selective Imputation for Incomplete Multi-view Clustering
von: Xu, Cai, et al.
Veröffentlicht: (2025)