Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Ye, Junyan, He, Jun, Li, Weijia, Lv, Zhutao, Lin, Yi, Yu, Jinhua, Yang, Haote, He, Conghui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Cross-view image geo-localization with Panorama-BEV Co-Retrieval Network
di: Ye, Junyan, et al.
Pubblicazione: (2024)
di: Ye, Junyan, et al.
Pubblicazione: (2024)
SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
di: Ye, Junyan, et al.
Pubblicazione: (2024)
di: Ye, Junyan, et al.
Pubblicazione: (2024)
Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
di: Feng, Peilin, et al.
Pubblicazione: (2025)
di: Feng, Peilin, et al.
Pubblicazione: (2025)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
di: Zhou, Baichuan, et al.
Pubblicazione: (2024)
di: Zhou, Baichuan, et al.
Pubblicazione: (2024)
3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level Supervisions
di: Li, Weijia, et al.
Pubblicazione: (2024)
di: Li, Weijia, et al.
Pubblicazione: (2024)
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
di: Zhu, Leqi, et al.
Pubblicazione: (2026)
di: Zhu, Leqi, et al.
Pubblicazione: (2026)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
di: Li, Weijia, et al.
Pubblicazione: (2024)
di: Li, Weijia, et al.
Pubblicazione: (2024)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
di: Yan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Yan, Zhiyuan, et al.
Pubblicazione: (2025)
LEGION: Learning to Ground and Explain for Synthetic Image Detection
di: Kang, Hengrui, et al.
Pubblicazione: (2025)
di: Kang, Hengrui, et al.
Pubblicazione: (2025)
OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild
di: Guo, Yuncheng, et al.
Pubblicazione: (2025)
di: Guo, Yuncheng, et al.
Pubblicazione: (2025)
Fine-Grained Building Function Recognition from Street-View Images via Geometry-Aware Semi-Supervised Learning
di: Li, Weijia, et al.
Pubblicazione: (2024)
di: Li, Weijia, et al.
Pubblicazione: (2024)
BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
GenClaw: Code-Driven Agentic Image Generation
di: Ye, Junyan, et al.
Pubblicazione: (2026)
di: Ye, Junyan, et al.
Pubblicazione: (2026)
Where am I? Cross-View Geo-localization with Natural Language Descriptions
di: Ye, Junyan, et al.
Pubblicazione: (2024)
di: Ye, Junyan, et al.
Pubblicazione: (2024)
Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse Exploration
di: Huang, Zilong, et al.
Pubblicazione: (2025)
di: Huang, Zilong, et al.
Pubblicazione: (2025)
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation
di: Wen, Siwei, et al.
Pubblicazione: (2025)
di: Wen, Siwei, et al.
Pubblicazione: (2025)
UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective
di: He, Jun, et al.
Pubblicazione: (2025)
di: He, Jun, et al.
Pubblicazione: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
di: He, Jun, et al.
Pubblicazione: (2026)
di: He, Jun, et al.
Pubblicazione: (2026)
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
di: Ye, Junyan, et al.
Pubblicazione: (2024)
di: Ye, Junyan, et al.
Pubblicazione: (2024)
BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving
di: Brandstaetter, Felix, et al.
Pubblicazione: (2025)
di: Brandstaetter, Felix, et al.
Pubblicazione: (2025)
MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts
di: Huang, Zilong, et al.
Pubblicazione: (2025)
di: Huang, Zilong, et al.
Pubblicazione: (2025)
BEV-Patch-PF: Particle Filtering with BEV-Aerial Feature Matching for Off-Road Geo-Localization
di: Lee, Dongmyeong, et al.
Pubblicazione: (2025)
di: Lee, Dongmyeong, et al.
Pubblicazione: (2025)
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
di: Vuong, Khiem, et al.
Pubblicazione: (2025)
di: Vuong, Khiem, et al.
Pubblicazione: (2025)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
di: Wen, Zichen, et al.
Pubblicazione: (2025)
di: Wen, Zichen, et al.
Pubblicazione: (2025)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
di: Gao, Junyuan, et al.
Pubblicazione: (2025)
di: Gao, Junyuan, et al.
Pubblicazione: (2025)
VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
di: Pang, Chao, et al.
Pubblicazione: (2024)
di: Pang, Chao, et al.
Pubblicazione: (2024)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
di: Kang, Hengrui, et al.
Pubblicazione: (2025)
di: Kang, Hengrui, et al.
Pubblicazione: (2025)
CLIP-BEVFormer: Enhancing Multi-View Image-Based BEV Detector with Ground Truth Flow
di: Pan, Chenbin, et al.
Pubblicazione: (2024)
di: Pan, Chenbin, et al.
Pubblicazione: (2024)
End-to-End Driving with Online Trajectory Evaluation via BEV World Model
di: Li, Yingyan, et al.
Pubblicazione: (2025)
di: Li, Yingyan, et al.
Pubblicazione: (2025)
Parrot Captions Teach CLIP to Spot Text
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
di: Lin, Honglin, et al.
Pubblicazione: (2026)
di: Lin, Honglin, et al.
Pubblicazione: (2026)
GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
di: Song, Ziying, et al.
Pubblicazione: (2024)
di: Song, Ziying, et al.
Pubblicazione: (2024)
Skyeyes: Ground Roaming using Aerial View Images
di: Gao, Zhiyuan, et al.
Pubblicazione: (2024)
di: Gao, Zhiyuan, et al.
Pubblicazione: (2024)
SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors
di: Fan, Ruijie, et al.
Pubblicazione: (2025)
di: Fan, Ruijie, et al.
Pubblicazione: (2025)
LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification
di: Zhang, Pingping, et al.
Pubblicazione: (2025)
di: Zhang, Pingping, et al.
Pubblicazione: (2025)
EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation
di: Li, Zhijiang, et al.
Pubblicazione: (2025)
di: Li, Zhijiang, et al.
Pubblicazione: (2025)
MaskBEV: Towards A Unified Framework for BEV Detection and Map Segmentation
di: Zhao, Xiao, et al.
Pubblicazione: (2024)
di: Zhao, Xiao, et al.
Pubblicazione: (2024)
Hierarchical and Decoupled BEV Perception Learning Framework for Autonomous Driving
di: Dai, Yuqi, et al.
Pubblicazione: (2024)
di: Dai, Yuqi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Cross-view image geo-localization with Panorama-BEV Co-Retrieval Network
di: Ye, Junyan, et al.
Pubblicazione: (2024) -
SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
di: Ye, Junyan, et al.
Pubblicazione: (2024) -
Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
di: Feng, Peilin, et al.
Pubblicazione: (2025) -
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
di: Zhou, Baichuan, et al.
Pubblicazione: (2024) -
3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level Supervisions
di: Li, Weijia, et al.
Pubblicazione: (2024)