PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Wenmiao, Zhang, Yichen, Liang, Yuxuan, Han, Xianjing, Yin, Yifang, Kruppa, Hannes, Ng, See-Kiong, Zimmermann, Roger |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance
von: Liu, Renyang, et al.
Veröffentlicht: (2024)
von: Liu, Renyang, et al.
Veröffentlicht: (2024)
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
von: Chen, Shengkai, et al.
Veröffentlicht: (2025)
von: Chen, Shengkai, et al.
Veröffentlicht: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
Robust Fuzzy Multi-view Learning under View Conflict
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
Multi-view Hypergraph-based Contrastive Learning Model for Cold-Start Micro-video Recommendation
von: Lyu, Sisuo, et al.
Veröffentlicht: (2024)
von: Lyu, Sisuo, et al.
Veröffentlicht: (2024)
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Calibrated Multimodal Representation Learning with Missing Modalities
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
High-level Codes and Fine-grained Weights for Online Multi-modal Hashing Retrieval
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
von: Xie, Liping, et al.
Veröffentlicht: (2025)
von: Xie, Liping, et al.
Veröffentlicht: (2025)
Simple Yet Effective Selective Imputation for Incomplete Multi-view Clustering
von: Xu, Cai, et al.
Veröffentlicht: (2025)
von: Xu, Cai, et al.
Veröffentlicht: (2025)
Deep Contrastive Multi-view Clustering under Semantic Feature Guidance
von: Liu, Siwen, et al.
Veröffentlicht: (2024)
von: Liu, Siwen, et al.
Veröffentlicht: (2024)
Regularized Contrastive Partial Multi-view Outlier Detection
von: Wang, Yijia, et al.
Veröffentlicht: (2024)
von: Wang, Yijia, et al.
Veröffentlicht: (2024)
MarsSQE: Stereo Quality Enhancement for Martian Images Using Bi-level Cross-view Attention
von: Xu, Mai, et al.
Veröffentlicht: (2024)
von: Xu, Mai, et al.
Veröffentlicht: (2024)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2025)
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2025)
Nutrition Estimation for Dietary Management: A Transformer Approach with Depth Sensing
von: Kwan, Zhengyi, et al.
Veröffentlicht: (2024)
von: Kwan, Zhengyi, et al.
Veröffentlicht: (2024)
MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Short-Form Video Viewing Behavior Analysis and Multi-Step Viewing Time Prediction
von: Yen, Vu Thi Hai, et al.
Veröffentlicht: (2026)
von: Yen, Vu Thi Hai, et al.
Veröffentlicht: (2026)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
altiro3D: Scene representation from single image and novel view synthesis
von: Canessa, E., et al.
Veröffentlicht: (2023)
von: Canessa, E., et al.
Veröffentlicht: (2023)
VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming with Edge Computing
von: Hu, Qiang, et al.
Veröffentlicht: (2025)
von: Hu, Qiang, et al.
Veröffentlicht: (2025)
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
von: He, Xusheng, et al.
Veröffentlicht: (2025)
von: He, Xusheng, et al.
Veröffentlicht: (2025)
RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
SpeechEE: A Novel Benchmark for Speech Event Extraction
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
TOL: Textual Localization with OpenStreetMap
von: Liao, Youqi, et al.
Veröffentlicht: (2026)
von: Liao, Youqi, et al.
Veröffentlicht: (2026)
Volume Tracking Based Reference Mesh Extraction for Time-Varying Mesh Compression
von: Chen, Guodong, et al.
Veröffentlicht: (2024)
von: Chen, Guodong, et al.
Veröffentlicht: (2024)
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2025)
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance
von: Liu, Renyang, et al.
Veröffentlicht: (2024) -
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025) -
OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
von: Chen, Shengkai, et al.
Veröffentlicht: (2025) -
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026) -
Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
von: Bin, Yi, et al.
Veröffentlicht: (2024)