Robust Visual Localization via Semantic-Guided Multi-Scale Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Zhongtao, Huang, Wenhao, Chen, Zhidong, Sun, Xiao Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
von: Ren, Li, et al.
Veröffentlicht: (2025)
von: Ren, Li, et al.
Veröffentlicht: (2025)
GESS: Multi-cue Guided Local Feature Learning via Geometric and Semantic Synergy
von: Yi, Yang, et al.
Veröffentlicht: (2026)
von: Yi, Yang, et al.
Veröffentlicht: (2026)
MultiLoc: Multi-view Guided Relative Pose Regression for Fast and Robust Visual Re-Localization
von: Dang, Nobel, et al.
Veröffentlicht: (2026)
von: Dang, Nobel, et al.
Veröffentlicht: (2026)
MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping
von: Fateh, Amirreza, et al.
Veröffentlicht: (2024)
von: Fateh, Amirreza, et al.
Veröffentlicht: (2024)
Semantic Visual Simultaneous Localization and Mapping: A Survey
von: Chen, Kaiqi, et al.
Veröffentlicht: (2022)
von: Chen, Kaiqi, et al.
Veröffentlicht: (2022)
Towards Scale-Aware Low-Light Enhancement via Structure-Guided Transformer Design
von: Dong, Wei, et al.
Veröffentlicht: (2025)
von: Dong, Wei, et al.
Veröffentlicht: (2025)
BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion
von: Tian, Qingyao, et al.
Veröffentlicht: (2026)
von: Tian, Qingyao, et al.
Veröffentlicht: (2026)
TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting
von: Liu, Taorong, et al.
Veröffentlicht: (2023)
von: Liu, Taorong, et al.
Veröffentlicht: (2023)
Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
von: Huang, Xian-Hong, et al.
Veröffentlicht: (2025)
von: Huang, Xian-Hong, et al.
Veröffentlicht: (2025)
Image Forgery Localization via Guided Noise and Multi-Scale Feature Aggregation
von: Niu, Yakun, et al.
Veröffentlicht: (2024)
von: Niu, Yakun, et al.
Veröffentlicht: (2024)
RealisID: Scale-Robust and Fine-Controllable Identity Customization via Local and Global Complementation
von: Sun, Zhaoyang, et al.
Veröffentlicht: (2024)
von: Sun, Zhaoyang, et al.
Veröffentlicht: (2024)
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception
von: Chen, Chuang, et al.
Veröffentlicht: (2024)
von: Chen, Chuang, et al.
Veröffentlicht: (2024)
Visually-Guided Controllable Medical Image Generation via Fine-Grained Semantic Disentanglement
von: Huang, Xin, et al.
Veröffentlicht: (2026)
von: Huang, Xin, et al.
Veröffentlicht: (2026)
Semantic and Feature Guided Uncertainty Quantification of Visual Localization for Autonomous Vehicles
von: Wu, Qiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Qiyuan, et al.
Veröffentlicht: (2025)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
von: Yin, Wei, et al.
Veröffentlicht: (2022)
von: Yin, Wei, et al.
Veröffentlicht: (2022)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
von: Fan, Bing, et al.
Veröffentlicht: (2025)
von: Fan, Bing, et al.
Veröffentlicht: (2025)
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
A Transformer-Based Adaptive Semantic Aggregation Method for UAV Visual Geo-Localization
von: Li, Shishen, et al.
Veröffentlicht: (2024)
von: Li, Shishen, et al.
Veröffentlicht: (2024)
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
von: Liang, Susan, et al.
Veröffentlicht: (2024)
von: Liang, Susan, et al.
Veröffentlicht: (2024)
SignNav: Leveraging Signage for Semantic Visual Navigation in Large-Scale Indoor Environments
von: Sun, Jian, et al.
Veröffentlicht: (2026)
von: Sun, Jian, et al.
Veröffentlicht: (2026)
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
IRIS-SLAM: Unified Geo-Instance Representations for Robust Semantic Localization and Mapping
von: Xiao, Tingyang, et al.
Veröffentlicht: (2026)
von: Xiao, Tingyang, et al.
Veröffentlicht: (2026)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
Semantic Localization Guiding Segment Anything Model For Reference Remote Sensing Image Segmentation
von: Li, Shuyang, et al.
Veröffentlicht: (2025)
von: Li, Shuyang, et al.
Veröffentlicht: (2025)
MIGA: Mutual Information-Guided Attack on Denoising Models for Semantic Manipulation
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
Morphology-optimized Multi-Scale Fusion: Combining Local Artifacts and Mesoscopic Semantics for Deepfake Detection and Localization
von: Shuai, Chao, et al.
Veröffentlicht: (2025)
von: Shuai, Chao, et al.
Veröffentlicht: (2025)
View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
MSGL-Transformer: A Multi-Scale Global-Local Transformer for Rodent Social Behavior Recognition
von: Sharif, Muhammad Imran, et al.
Veröffentlicht: (2026)
von: Sharif, Muhammad Imran, et al.
Veröffentlicht: (2026)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
von: Park, Chunghyun, et al.
Veröffentlicht: (2024)
von: Park, Chunghyun, et al.
Veröffentlicht: (2024)
Boosting Adversarial Transferability Against Defenses via Multi-Scale Transformation
von: Guo, Zihong, et al.
Veröffentlicht: (2025)
von: Guo, Zihong, et al.
Veröffentlicht: (2025)
Semantic-Guided Global-Local Collaborative Networks for Lightweight Image Super-Resolution
von: Fan, Wanshu, et al.
Veröffentlicht: (2025)
von: Fan, Wanshu, et al.
Veröffentlicht: (2025)
LACV-Net: Semantic Segmentation of Large-Scale Point Cloud Scene via Local Adaptive and Comprehensive VLAD
von: Zeng, Ziyin, et al.
Veröffentlicht: (2022)
von: Zeng, Ziyin, et al.
Veröffentlicht: (2022)
Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
von: Peng, Zhong, et al.
Veröffentlicht: (2025)
von: Peng, Zhong, et al.
Veröffentlicht: (2025)
R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization
von: Jiang, Xudong, et al.
Veröffentlicht: (2025)
von: Jiang, Xudong, et al.
Veröffentlicht: (2025)
Semantic Guided Large Scale Factor Remote Sensing Image Super-resolution with Generative Diffusion Prior
von: Wang, Ce, et al.
Veröffentlicht: (2024)
von: Wang, Ce, et al.
Veröffentlicht: (2024)
Multi-dimensional Visual Prompt Enhanced Image Restoration via Mamba-Transformer Aggregation
von: Jiang, Aiwen, et al.
Veröffentlicht: (2024)
von: Jiang, Aiwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
von: Ren, Li, et al.
Veröffentlicht: (2025) -
GESS: Multi-cue Guided Local Feature Learning via Geometric and Semantic Synergy
von: Yi, Yang, et al.
Veröffentlicht: (2026) -
MultiLoc: Multi-view Guided Relative Pose Regression for Fast and Robust Visual Re-Localization
von: Dang, Nobel, et al.
Veröffentlicht: (2026) -
MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping
von: Fateh, Amirreza, et al.
Veröffentlicht: (2024) -
Semantic Visual Simultaneous Localization and Mapping: A Survey
von: Chen, Kaiqi, et al.
Veröffentlicht: (2022)