GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Hongyang, Liu, Yinhao, Zhang, Haitao, Wen, Zhongyi, Kuang, Zhenyu, Liang, Shuxian, Hua, Xiansheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SkyLink: Unifying Street-Satellite Geo-Localization via UAV-Mediated 3D Scene Alignment
di: Zhang, Hongyang, et al.
Pubblicazione: (2025)
di: Zhang, Hongyang, et al.
Pubblicazione: (2025)
Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis
di: Dong, Guangyuan, et al.
Pubblicazione: (2026)
di: Dong, Guangyuan, et al.
Pubblicazione: (2026)
Dual Dynamic Threshold Adjustment Strategy for Deep Metric Learning
di: Jiang, Xiruo, et al.
Pubblicazione: (2024)
di: Jiang, Xiruo, et al.
Pubblicazione: (2024)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
di: Li, Qingcao, et al.
Pubblicazione: (2026)
di: Li, Qingcao, et al.
Pubblicazione: (2026)
OpenVNA: A Framework for Analyzing the Behavior of Multimodal Language Understanding System under Noisy Scenarios
di: Yuan, Ziqi, et al.
Pubblicazione: (2024)
di: Yuan, Ziqi, et al.
Pubblicazione: (2024)
PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology
di: Sun, Yuxuan, et al.
Pubblicazione: (2023)
di: Sun, Yuxuan, et al.
Pubblicazione: (2023)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
di: Mao, Junzhu, et al.
Pubblicazione: (2025)
di: Mao, Junzhu, et al.
Pubblicazione: (2025)
Contrastive Knowledge Distillation for Robust Multimodal Sentiment Analysis
di: Sang, Zhongyi, et al.
Pubblicazione: (2024)
di: Sang, Zhongyi, et al.
Pubblicazione: (2024)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
di: Hu, Wenmiao, et al.
Pubblicazione: (2024)
di: Hu, Wenmiao, et al.
Pubblicazione: (2024)
High-Fidelity 3D Gaussian Human Reconstruction via Region-Aware Initialization and Geometric Priors
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
di: Xu, Yang, et al.
Pubblicazione: (2024)
di: Xu, Yang, et al.
Pubblicazione: (2024)
PiGW: A Plug-in Generative Watermarking Framework
di: Ma, Rui, et al.
Pubblicazione: (2024)
di: Ma, Rui, et al.
Pubblicazione: (2024)
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
di: Fan, Congyi, et al.
Pubblicazione: (2026)
di: Fan, Congyi, et al.
Pubblicazione: (2026)
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
di: Tong, Haonan, et al.
Pubblicazione: (2024)
di: Tong, Haonan, et al.
Pubblicazione: (2024)
VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming with Edge Computing
di: Hu, Qiang, et al.
Pubblicazione: (2025)
di: Hu, Qiang, et al.
Pubblicazione: (2025)
Smaller is Better: Generative Models Can Power Short Video Preloading
di: Liu, Liming, et al.
Pubblicazione: (2026)
di: Liu, Liming, et al.
Pubblicazione: (2026)
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
di: Yin, Jianjian, et al.
Pubblicazione: (2025)
di: Yin, Jianjian, et al.
Pubblicazione: (2025)
Towards Better Text-to-Image Generation Alignment via Attention Modulation
di: Wu, Yihang, et al.
Pubblicazione: (2024)
di: Wu, Yihang, et al.
Pubblicazione: (2024)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
di: Wang, Siyu, et al.
Pubblicazione: (2025)
di: Wang, Siyu, et al.
Pubblicazione: (2025)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
di: Cheng, Fenghua, et al.
Pubblicazione: (2025)
di: Cheng, Fenghua, et al.
Pubblicazione: (2025)
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning
di: Hu, Jiayun, et al.
Pubblicazione: (2025)
di: Hu, Jiayun, et al.
Pubblicazione: (2025)
A 3D-Cascading Crossing Coupling Framework for Hyperchaotic Map Construction and Its Application to Color Image Encryption
di: Sun, Jilei, et al.
Pubblicazione: (2025)
di: Sun, Jilei, et al.
Pubblicazione: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
MViR: Multi-View Visual-Semantic Representation for Fake News Detection
di: Liang, Haochen, et al.
Pubblicazione: (2026)
di: Liang, Haochen, et al.
Pubblicazione: (2026)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
di: Zhou, Hengyang, et al.
Pubblicazione: (2025)
di: Zhou, Hengyang, et al.
Pubblicazione: (2025)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
di: Zhou, Qianrui, et al.
Pubblicazione: (2023)
di: Zhou, Qianrui, et al.
Pubblicazione: (2023)
Quality-Aware Dynamic Resolution Adaptation Framework for Adaptive Video Streaming
di: Premkumar, Amritha, et al.
Pubblicazione: (2024)
di: Premkumar, Amritha, et al.
Pubblicazione: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
di: Zhang, Han, et al.
Pubblicazione: (2025)
di: Zhang, Han, et al.
Pubblicazione: (2025)
Cross-Platform Neural Video Coding: A Case Study
di: Conceição, Ruhan, et al.
Pubblicazione: (2024)
di: Conceição, Ruhan, et al.
Pubblicazione: (2024)
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
di: Pan, Zhaoyan, et al.
Pubblicazione: (2026)
di: Pan, Zhaoyan, et al.
Pubblicazione: (2026)
Short-Form Video Viewing Behavior Analysis and Multi-Step Viewing Time Prediction
di: Yen, Vu Thi Hai, et al.
Pubblicazione: (2026)
di: Yen, Vu Thi Hai, et al.
Pubblicazione: (2026)
Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition
di: Li, Qilin, et al.
Pubblicazione: (2025)
di: Li, Qilin, et al.
Pubblicazione: (2025)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
di: Huang, Yiheng, et al.
Pubblicazione: (2024)
di: Huang, Yiheng, et al.
Pubblicazione: (2024)
Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
di: Zhao, Xianbing, et al.
Pubblicazione: (2025)
di: Zhao, Xianbing, et al.
Pubblicazione: (2025)
Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition
di: Zhuang, Yan, et al.
Pubblicazione: (2026)
di: Zhuang, Yan, et al.
Pubblicazione: (2026)
An Emotion Recognition Framework via Cross-modal Alignment of EEG and Eye Movement Data
di: Wang, Jianlu, et al.
Pubblicazione: (2025)
di: Wang, Jianlu, et al.
Pubblicazione: (2025)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
di: Qian, Zhiwen, et al.
Pubblicazione: (2025)
di: Qian, Zhiwen, et al.
Pubblicazione: (2025)
LLM2Manim: Pedagogy-Aware AI Generation of STEM Animations
di: Joshi, Aastha, et al.
Pubblicazione: (2026)
di: Joshi, Aastha, et al.
Pubblicazione: (2026)
Period-conscious Time-series Reconstruction under Local Differential Privacy
di: Wang, Yaxuan, et al.
Pubblicazione: (2026)
di: Wang, Yaxuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SkyLink: Unifying Street-Satellite Geo-Localization via UAV-Mediated 3D Scene Alignment
di: Zhang, Hongyang, et al.
Pubblicazione: (2025) -
Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis
di: Dong, Guangyuan, et al.
Pubblicazione: (2026) -
Dual Dynamic Threshold Adjustment Strategy for Deep Metric Learning
di: Jiang, Xiruo, et al.
Pubblicazione: (2024) -
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
di: Li, Qingcao, et al.
Pubblicazione: (2026) -
OpenVNA: A Framework for Analyzing the Behavior of Multimodal Language Understanding System under Noisy Scenarios
di: Yuan, Ziqi, et al.
Pubblicazione: (2024)