From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Bajbaa, Khawlah, Anwar, Abbas, Saqib, Muhammad, Anwar, Hafeez, Sharma, Nabin, Usman, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bird Eye-View to Street-View: A Survey
by: Bajbaa, Khawlah, et al.
Published: (2024)
by: Bajbaa, Khawlah, et al.
Published: (2024)
PanoGAN A Deep Generative Model for Panoramic Dental Radiographs
by: Pedersen, Soren, et al.
Published: (2025)
by: Pedersen, Soren, et al.
Published: (2025)
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
by: Xie, Tianyidan, et al.
Published: (2026)
by: Xie, Tianyidan, et al.
Published: (2026)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
by: Hu, Wenmiao, et al.
Published: (2024)
by: Hu, Wenmiao, et al.
Published: (2024)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
by: Cheng, Fenghua, et al.
Published: (2025)
by: Cheng, Fenghua, et al.
Published: (2025)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
by: Yang, Quanwei, et al.
Published: (2025)
by: Yang, Quanwei, et al.
Published: (2025)
NeRF View Synthesis: Subjective Quality Assessment and Objective Metrics Evaluation
by: Martin, Pedro, et al.
Published: (2024)
by: Martin, Pedro, et al.
Published: (2024)
MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection
by: Alghamdi, Leena, et al.
Published: (2025)
by: Alghamdi, Leena, et al.
Published: (2025)
SkyLink: Unifying Street-Satellite Geo-Localization via UAV-Mediated 3D Scene Alignment
by: Zhang, Hongyang, et al.
Published: (2025)
by: Zhang, Hongyang, et al.
Published: (2025)
Intelligent Carrier Allocation: A Cross-Modal Reasoning Framework for Adaptive Multimodal Steganography
by: Das, Abhirup, et al.
Published: (2025)
by: Das, Abhirup, et al.
Published: (2025)
Cardiverse: Harnessing LLMs for Novel Card Game Prototyping
by: Li, Danrui, et al.
Published: (2025)
by: Li, Danrui, et al.
Published: (2025)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
by: Tian, Chong, et al.
Published: (2026)
by: Tian, Chong, et al.
Published: (2026)
Short-Form Video Viewing Behavior Analysis and Multi-Step Viewing Time Prediction
by: Yen, Vu Thi Hai, et al.
Published: (2026)
by: Yen, Vu Thi Hai, et al.
Published: (2026)
GANonymization: A GAN-based Face Anonymization Framework for Preserving Emotional Expressions
by: Hellmann, Fabio, et al.
Published: (2023)
by: Hellmann, Fabio, et al.
Published: (2023)
VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming with Edge Computing
by: Hu, Qiang, et al.
Published: (2025)
by: Hu, Qiang, et al.
Published: (2025)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
An Emotion Recognition Framework via Cross-modal Alignment of EEG and Eye Movement Data
by: Wang, Jianlu, et al.
Published: (2025)
by: Wang, Jianlu, et al.
Published: (2025)
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
by: Fan, Congyi, et al.
Published: (2026)
by: Fan, Congyi, et al.
Published: (2026)
RoSMM: A Robust and Secure Multi-Modal Watermarking Framework for Diffusion Models
by: Fang, ZhongLi, et al.
Published: (2025)
by: Fang, ZhongLi, et al.
Published: (2025)
DiffCL: A Diffusion-Based Contrastive Learning Framework with Semantic Alignment for Multimodal Recommendations
by: Song, Qiya, et al.
Published: (2025)
by: Song, Qiya, et al.
Published: (2025)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
by: Huang, Yiheng, et al.
Published: (2024)
by: Huang, Yiheng, et al.
Published: (2024)
GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization
by: Zhang, Hongyang, et al.
Published: (2026)
by: Zhang, Hongyang, et al.
Published: (2026)
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
by: Yin, Yongkang, et al.
Published: (2023)
by: Yin, Yongkang, et al.
Published: (2023)
Deep Reversible Consistency Learning for Cross-modal Retrieval
by: Pu, Ruitao, et al.
Published: (2025)
by: Pu, Ruitao, et al.
Published: (2025)
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
A 3D-Cascading Crossing Coupling Framework for Hyperchaotic Map Construction and Its Application to Color Image Encryption
by: Sun, Jilei, et al.
Published: (2025)
by: Sun, Jilei, et al.
Published: (2025)
Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency
by: Zarghani, Abolfazl, et al.
Published: (2025)
by: Zarghani, Abolfazl, et al.
Published: (2025)
Real-Time Interactive Hybrid Ocean: Spectrum-Consistent Wave Particle-FFT Coupling
by: Xue, Shengze, et al.
Published: (2025)
by: Xue, Shengze, et al.
Published: (2025)
Integration of Policy and Reputation based Trust Mechanisms in e-Commerce Industry
by: Siddiqui, Muhammad Yasir, et al.
Published: (2024)
by: Siddiqui, Muhammad Yasir, et al.
Published: (2024)
TOL: Textual Localization with OpenStreetMap
by: Liao, Youqi, et al.
Published: (2026)
by: Liao, Youqi, et al.
Published: (2026)
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
by: Xiao, Junhao, et al.
Published: (2025)
by: Xiao, Junhao, et al.
Published: (2025)
Fact-Checking with Contextual Narratives: Leveraging Retrieval-Augmented LLMs for Social Media Analysis
by: Dey, Arka Ujjal, et al.
Published: (2025)
by: Dey, Arka Ujjal, et al.
Published: (2025)
CDI-DTI: A Strong Cross-domain Interpretable Drug-Target Interaction Prediction Framework Based on Multi-Strategy Fusion
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
HyDiscGAN: A Hybrid Distributed cGAN for Audio-Visual Privacy Preservation in Multimodal Sentiment Analysis
by: Wu, Zhuojia, et al.
Published: (2024)
by: Wu, Zhuojia, et al.
Published: (2024)
Real-Time Position-Aware View Synthesis from Single-View Input
by: Gond, Manu, et al.
Published: (2024)
by: Gond, Manu, et al.
Published: (2024)
MViR: Multi-View Visual-Semantic Representation for Fake News Detection
by: Liang, Haochen, et al.
Published: (2026)
by: Liang, Haochen, et al.
Published: (2026)
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
by: Pan, Zhaoyan, et al.
Published: (2026)
by: Pan, Zhaoyan, et al.
Published: (2026)
Disparity-based Stereo Image Compression with Aligned Cross-View Priors
by: Zhai, Yongqi, et al.
Published: (2022)
by: Zhai, Yongqi, et al.
Published: (2022)
BOLA360: Near-optimal View and Bitrate Adaptation for 360-degree Video Streaming
by: Zeynali, Ali, et al.
Published: (2023)
by: Zeynali, Ali, et al.
Published: (2023)
Verifying Cross-modal Entity Consistency in News using Vision-language Models
by: Tahmasebi, Sahar, et al.
Published: (2025)
by: Tahmasebi, Sahar, et al.
Published: (2025)
Similar Items
-
Bird Eye-View to Street-View: A Survey
by: Bajbaa, Khawlah, et al.
Published: (2024) -
PanoGAN A Deep Generative Model for Panoramic Dental Radiographs
by: Pedersen, Soren, et al.
Published: (2025) -
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
by: Xie, Tianyidan, et al.
Published: (2026) -
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
by: Hu, Wenmiao, et al.
Published: (2024) -
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
by: Cheng, Fenghua, et al.
Published: (2025)