Bridging Text and Vision: A Multi-View Text-Vision Registration Approach for Cross-Modal Place Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Shang, Tianyi, Li, Zhenyu, Xu, Pengjie, Qiao, Jinwei, Chen, Gang, Ruan, Zihan, Hu, Weijun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MambaPlace:Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba Mechanisms
by: Shang, Tianyi, et al.
Published: (2024)
by: Shang, Tianyi, et al.
Published: (2024)
Vehicle-Scene Interaction: A Text-Driven 3D Lidar Place Recognition Method for Autonomous Driving
by: Shang, Tianyi, et al.
Published: (2025)
by: Shang, Tianyi, et al.
Published: (2025)
Riemannian and Symplectic Geometry for Hierarchical Text-Driven Place Recognition
by: Shang, Tianyi, et al.
Published: (2026)
by: Shang, Tianyi, et al.
Published: (2026)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
by: Tao, Huaqi, et al.
Published: (2025)
by: Tao, Huaqi, et al.
Published: (2025)
Place Recognition Meet Multiple Modalitie: A Comprehensive Review, Current Challenges and Future Directions
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
by: Xu, Zhengyi, et al.
Published: (2026)
by: Xu, Zhengyi, et al.
Published: (2026)
Monocular Visual Place Recognition in LiDAR Maps via Cross-Modal State Space Model and Multi-View Matching
by: Yao, Gongxin, et al.
Published: (2024)
by: Yao, Gongxin, et al.
Published: (2024)
OptiCorNet: Optimizing Sequence-Based Context Correlation for Visual Place Recognition
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation
by: Hu, Zixuan, et al.
Published: (2026)
by: Hu, Zixuan, et al.
Published: (2026)
SpatiaLoc: Leveraging Multi-Level Spatial Enhanced Descriptors for Cross-Modal Localization
by: Shang, Tianyi, et al.
Published: (2026)
by: Shang, Tianyi, et al.
Published: (2026)
DiffPlace: Street View Generation via Place-Controllable Diffusion Model Enhancing Place Recognition
by: Li, Ji, et al.
Published: (2026)
by: Li, Ji, et al.
Published: (2026)
Systematic Evaluation of Novel View Synthesis for Video Place Recognition
by: Mahmud, Muhammad Zawad, et al.
Published: (2026)
by: Mahmud, Muhammad Zawad, et al.
Published: (2026)
DisPlace: Discriminative Place Projections for Multi-Reference Visual Place Recognition
by: Rajani, Dhyey Manish, et al.
Published: (2026)
by: Rajani, Dhyey Manish, et al.
Published: (2026)
Pair-VPR: Place-Aware Pre-training and Contrastive Pair Classification for Visual Place Recognition with Vision Transformers
by: Hausler, Stephen, et al.
Published: (2024)
by: Hausler, Stephen, et al.
Published: (2024)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
by: Li, Haosheng, et al.
Published: (2026)
by: Li, Haosheng, et al.
Published: (2026)
GOTPR: General Outdoor Text-based Place Recognition Using Scene Graph Retrieval with OpenStreetMap
by: Jung, Donghwi, et al.
Published: (2025)
by: Jung, Donghwi, et al.
Published: (2025)
ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
by: Xie, Weidong, et al.
Published: (2024)
by: Xie, Weidong, et al.
Published: (2024)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
CSCPR: Cross-Source-Context Indoor RGB-D Place Recognition
by: Liang, Jing, et al.
Published: (2024)
by: Liang, Jing, et al.
Published: (2024)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
CLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models
by: Pueyo, Pablo, et al.
Published: (2024)
by: Pueyo, Pablo, et al.
Published: (2024)
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition
by: Xiong, Songsong, et al.
Published: (2025)
by: Xiong, Songsong, et al.
Published: (2025)
BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View Images
by: Luo, Lun, et al.
Published: (2023)
by: Luo, Lun, et al.
Published: (2023)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities
by: Zhou, Yiyun, et al.
Published: (2025)
by: Zhou, Yiyun, et al.
Published: (2025)
RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception
by: Ma, Jiahao, et al.
Published: (2026)
by: Ma, Jiahao, et al.
Published: (2026)
DSFormer: A Dual-Scale Cross-Learning Transformer for Visual Place Recognition
by: Jiang, Haiyang, et al.
Published: (2025)
by: Jiang, Haiyang, et al.
Published: (2025)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
by: Liu, Youzhi, et al.
Published: (2024)
by: Liu, Youzhi, et al.
Published: (2024)
PlaceFormer: Transformer-based Visual Place Recognition using Multi-Scale Patch Selection and Fusion
by: Kannan, Shyam Sundar, et al.
Published: (2024)
by: Kannan, Shyam Sundar, et al.
Published: (2024)
CMR-Agent: Learning a Cross-Modal Agent for Iterative Image-to-Point Cloud Registration
by: Yao, Gongxin, et al.
Published: (2024)
by: Yao, Gongxin, et al.
Published: (2024)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
by: Huang, Huang, et al.
Published: (2025)
by: Huang, Huang, et al.
Published: (2025)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
Explicit Interaction for Fusion-Based Place Recognition
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial Views
by: Griffiths, Ethan, et al.
Published: (2025)
by: Griffiths, Ethan, et al.
Published: (2025)
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
VXP: Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition
by: Li, Yun-Jin, et al.
Published: (2024)
by: Li, Yun-Jin, et al.
Published: (2024)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
by: Holden, Lachlan, et al.
Published: (2026)
by: Holden, Lachlan, et al.
Published: (2026)
LCPR: A Multi-Scale Attention-Based LiDAR-Camera Fusion Network for Place Recognition
by: Zhou, Zijie, et al.
Published: (2023)
by: Zhou, Zijie, et al.
Published: (2023)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
by: Guo, Jianing, et al.
Published: (2025)
by: Guo, Jianing, et al.
Published: (2025)
Similar Items
-
MambaPlace:Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba Mechanisms
by: Shang, Tianyi, et al.
Published: (2024) -
Vehicle-Scene Interaction: A Text-Driven 3D Lidar Place Recognition Method for Autonomous Driving
by: Shang, Tianyi, et al.
Published: (2025) -
Riemannian and Symplectic Geometry for Hierarchical Text-Driven Place Recognition
by: Shang, Tianyi, et al.
Published: (2026) -
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
by: Tao, Huaqi, et al.
Published: (2025) -
Place Recognition Meet Multiple Modalitie: A Comprehensive Review, Current Challenges and Future Directions
by: Li, Zhenyu, et al.
Published: (2025)