Regressing Transformers for Data-efficient Visual Place Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Leyva-Vallina, María, Strisciuglio, Nicola, Petkov, Nicolai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
by: Vaish, Puru, et al.
Published: (2024)
by: Vaish, Puru, et al.
Published: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025)
by: Berasi, Davide, et al.
Published: (2025)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025)
by: Cheng, Jintao, et al.
Published: (2025)
Event-Driven Neuromorphic Vision Enables Energy-Efficient Visual Place Recognition
by: Keime, Geoffroy, et al.
Published: (2026)
by: Keime, Geoffroy, et al.
Published: (2026)
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
by: Ou, Yiwei, et al.
Published: (2025)
by: Ou, Yiwei, et al.
Published: (2025)
Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments
by: Shao, Xingyu, et al.
Published: (2026)
by: Shao, Xingyu, et al.
Published: (2026)
Map-Relative Pose Regression for Visual Re-Localization
by: Chen, Shuai, et al.
Published: (2024)
by: Chen, Shuai, et al.
Published: (2024)
CART: Compositional Auto-Regressive Transformer for Image Generation
by: Roheda, Siddharth, et al.
Published: (2024)
by: Roheda, Siddharth, et al.
Published: (2024)
Contrastive Learning for Regression on Hyperspectral Data
by: Dhaini, Mohamad, et al.
Published: (2024)
by: Dhaini, Mohamad, et al.
Published: (2024)
GViT: Representing Images as Gaussians for Visual Recognition
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
Revisit Anything: Visual Place Recognition via Image Segment Retrieval
by: Garg, Kartik, et al.
Published: (2024)
by: Garg, Kartik, et al.
Published: (2024)
EDTformer: An Efficient Decoder Transformer for Visual Place Recognition
by: Jin, Tong, et al.
Published: (2024)
by: Jin, Tong, et al.
Published: (2024)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Grounding Descriptions in Images informs Zero-Shot Visual Recognition
by: Halbe, Shaunak, et al.
Published: (2024)
by: Halbe, Shaunak, et al.
Published: (2024)
PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
by: Yang, Chenhongyi, et al.
Published: (2024)
by: Yang, Chenhongyi, et al.
Published: (2024)
Synthesizing Realistic Data for Table Recognition
by: Hou, Qiyu, et al.
Published: (2024)
by: Hou, Qiyu, et al.
Published: (2024)
Do ImageNet-trained models learn shortcuts? The impact of frequency shortcuts on generalization
by: Wang, Shunxin, et al.
Published: (2025)
by: Wang, Shunxin, et al.
Published: (2025)
TABLET: Table Structure Recognition using Encoder-only Transformers
by: Hou, Qiyu, et al.
Published: (2025)
by: Hou, Qiyu, et al.
Published: (2025)
Spectral-Spatial Contrastive Learning Framework for Regression on Hyperspectral Data
by: Dhaini, Mohamad, et al.
Published: (2026)
by: Dhaini, Mohamad, et al.
Published: (2026)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
by: Liang, Zichen, et al.
Published: (2025)
by: Liang, Zichen, et al.
Published: (2025)
Surely Large Multimodal Models (Don't) Excel in Visual Species Recognition?
by: Liu, Tian, et al.
Published: (2025)
by: Liu, Tian, et al.
Published: (2025)
Efficient Visual Transformer by Learnable Token Merging
by: Wang, Yancheng, et al.
Published: (2024)
by: Wang, Yancheng, et al.
Published: (2024)
ConvMixFormer- A Resource-efficient Convolution Mixer for Transformer-based Dynamic Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2024)
by: Garg, Mallika, et al.
Published: (2024)
Data-Driven Hierarchical Open Set Recognition
by: Hannum, Andrew, et al.
Published: (2024)
by: Hannum, Andrew, et al.
Published: (2024)
StrideNET: Swin Transformer for Terrain Recognition with Dynamic Roughness Extraction
by: Shelare, Maitreya, et al.
Published: (2024)
by: Shelare, Maitreya, et al.
Published: (2024)
Design and Analysis of Efficient Attention in Transformers for Social Group Activity Recognition
by: Tamura, Masato
Published: (2024)
by: Tamura, Masato
Published: (2024)
Classification and regression of trajectories rendered as images via 2D Convolutional Neural Networks
by: Nicolai, Mariaclaudia, et al.
Published: (2024)
by: Nicolai, Mariaclaudia, et al.
Published: (2024)
Enhancing Tea Leaf Disease Recognition with Attention Mechanisms and Grad-CAM Visualization
by: Shikdar, Omar Faruq, et al.
Published: (2025)
by: Shikdar, Omar Faruq, et al.
Published: (2025)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
by: Zheng, Ying, et al.
Published: (2025)
by: Zheng, Ying, et al.
Published: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)
by: Lee, Youngwan, et al.
Published: (2024)
HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimation
by: Zheng, Hongwei, et al.
Published: (2025)
by: Zheng, Hongwei, et al.
Published: (2025)
TransformMix: Learning Transformation and Mixing Strategies from Data
by: Cheung, Tsz-Him, et al.
Published: (2024)
by: Cheung, Tsz-Him, et al.
Published: (2024)
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Efficient License Plate Recognition in Videos Using Visual Rhythm and Accumulative Line Analysis
by: Ribeiro, Victor Nascimento, et al.
Published: (2025)
by: Ribeiro, Victor Nascimento, et al.
Published: (2025)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
MentalBlackboard: Evaluating Spatial Visualization via Mathematical Transformations
by: Yilmaz, Nilay, et al.
Published: (2026)
by: Yilmaz, Nilay, et al.
Published: (2026)
Brain Hematoma Marker Recognition Using Multitask Learning: SwinTransformer and Swin-Unet
by: Hirata, Kodai, et al.
Published: (2025)
by: Hirata, Kodai, et al.
Published: (2025)
A Transformer Based Handwriting Recognition System Jointly Using Online and Offline Features
by: Lodh, Ayush, et al.
Published: (2025)
by: Lodh, Ayush, et al.
Published: (2025)
Tab2Visual: Overcoming Limited Data in Tabular Data Classification Using Deep Learning with Visual Representations
by: Mamdouh, Ahmed, et al.
Published: (2025)
by: Mamdouh, Ahmed, et al.
Published: (2025)
Similar Items
-
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
by: Vaish, Puru, et al.
Published: (2024) -
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025) -
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025) -
Event-Driven Neuromorphic Vision Enables Energy-Efficient Visual Place Recognition
by: Keime, Geoffroy, et al.
Published: (2026) -
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
by: Ou, Yiwei, et al.
Published: (2025)