Regressing Transformers for Data-efficient Visual Place Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Leyva-Vallina, María, Strisciuglio, Nicola, Petkov, Nicolai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
von: Vaish, Puru, et al.
Veröffentlicht: (2024)
von: Vaish, Puru, et al.
Veröffentlicht: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
Event-Driven Neuromorphic Vision Enables Energy-Efficient Visual Place Recognition
von: Keime, Geoffroy, et al.
Veröffentlicht: (2026)
von: Keime, Geoffroy, et al.
Veröffentlicht: (2026)
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
von: Ou, Yiwei, et al.
Veröffentlicht: (2025)
von: Ou, Yiwei, et al.
Veröffentlicht: (2025)
Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments
von: Shao, Xingyu, et al.
Veröffentlicht: (2026)
von: Shao, Xingyu, et al.
Veröffentlicht: (2026)
Map-Relative Pose Regression for Visual Re-Localization
von: Chen, Shuai, et al.
Veröffentlicht: (2024)
von: Chen, Shuai, et al.
Veröffentlicht: (2024)
CART: Compositional Auto-Regressive Transformer for Image Generation
von: Roheda, Siddharth, et al.
Veröffentlicht: (2024)
von: Roheda, Siddharth, et al.
Veröffentlicht: (2024)
Contrastive Learning for Regression on Hyperspectral Data
von: Dhaini, Mohamad, et al.
Veröffentlicht: (2024)
von: Dhaini, Mohamad, et al.
Veröffentlicht: (2024)
GViT: Representing Images as Gaussians for Visual Recognition
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
Revisit Anything: Visual Place Recognition via Image Segment Retrieval
von: Garg, Kartik, et al.
Veröffentlicht: (2024)
von: Garg, Kartik, et al.
Veröffentlicht: (2024)
EDTformer: An Efficient Decoder Transformer for Visual Place Recognition
von: Jin, Tong, et al.
Veröffentlicht: (2024)
von: Jin, Tong, et al.
Veröffentlicht: (2024)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
Grounding Descriptions in Images informs Zero-Shot Visual Recognition
von: Halbe, Shaunak, et al.
Veröffentlicht: (2024)
von: Halbe, Shaunak, et al.
Veröffentlicht: (2024)
PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
von: Yang, Chenhongyi, et al.
Veröffentlicht: (2024)
von: Yang, Chenhongyi, et al.
Veröffentlicht: (2024)
Synthesizing Realistic Data for Table Recognition
von: Hou, Qiyu, et al.
Veröffentlicht: (2024)
von: Hou, Qiyu, et al.
Veröffentlicht: (2024)
Do ImageNet-trained models learn shortcuts? The impact of frequency shortcuts on generalization
von: Wang, Shunxin, et al.
Veröffentlicht: (2025)
von: Wang, Shunxin, et al.
Veröffentlicht: (2025)
TABLET: Table Structure Recognition using Encoder-only Transformers
von: Hou, Qiyu, et al.
Veröffentlicht: (2025)
von: Hou, Qiyu, et al.
Veröffentlicht: (2025)
Spectral-Spatial Contrastive Learning Framework for Regression on Hyperspectral Data
von: Dhaini, Mohamad, et al.
Veröffentlicht: (2026)
von: Dhaini, Mohamad, et al.
Veröffentlicht: (2026)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
von: Liang, Zichen, et al.
Veröffentlicht: (2025)
von: Liang, Zichen, et al.
Veröffentlicht: (2025)
Surely Large Multimodal Models (Don't) Excel in Visual Species Recognition?
von: Liu, Tian, et al.
Veröffentlicht: (2025)
von: Liu, Tian, et al.
Veröffentlicht: (2025)
Efficient Visual Transformer by Learnable Token Merging
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
ConvMixFormer- A Resource-efficient Convolution Mixer for Transformer-based Dynamic Hand Gesture Recognition
von: Garg, Mallika, et al.
Veröffentlicht: (2024)
von: Garg, Mallika, et al.
Veröffentlicht: (2024)
Data-Driven Hierarchical Open Set Recognition
von: Hannum, Andrew, et al.
Veröffentlicht: (2024)
von: Hannum, Andrew, et al.
Veröffentlicht: (2024)
StrideNET: Swin Transformer for Terrain Recognition with Dynamic Roughness Extraction
von: Shelare, Maitreya, et al.
Veröffentlicht: (2024)
von: Shelare, Maitreya, et al.
Veröffentlicht: (2024)
Design and Analysis of Efficient Attention in Transformers for Social Group Activity Recognition
von: Tamura, Masato
Veröffentlicht: (2024)
von: Tamura, Masato
Veröffentlicht: (2024)
Classification and regression of trajectories rendered as images via 2D Convolutional Neural Networks
von: Nicolai, Mariaclaudia, et al.
Veröffentlicht: (2024)
von: Nicolai, Mariaclaudia, et al.
Veröffentlicht: (2024)
Enhancing Tea Leaf Disease Recognition with Attention Mechanisms and Grad-CAM Visualization
von: Shikdar, Omar Faruq, et al.
Veröffentlicht: (2025)
von: Shikdar, Omar Faruq, et al.
Veröffentlicht: (2025)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
von: Zheng, Ying, et al.
Veröffentlicht: (2025)
von: Zheng, Ying, et al.
Veröffentlicht: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
Visualizing the loss landscape of Self-supervised Vision Transformer
von: Lee, Youngwan, et al.
Veröffentlicht: (2024)
von: Lee, Youngwan, et al.
Veröffentlicht: (2024)
HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimation
von: Zheng, Hongwei, et al.
Veröffentlicht: (2025)
von: Zheng, Hongwei, et al.
Veröffentlicht: (2025)
TransformMix: Learning Transformation and Mixing Strategies from Data
von: Cheung, Tsz-Him, et al.
Veröffentlicht: (2024)
von: Cheung, Tsz-Him, et al.
Veröffentlicht: (2024)
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2024)
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2024)
Efficient License Plate Recognition in Videos Using Visual Rhythm and Accumulative Line Analysis
von: Ribeiro, Victor Nascimento, et al.
Veröffentlicht: (2025)
von: Ribeiro, Victor Nascimento, et al.
Veröffentlicht: (2025)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
MentalBlackboard: Evaluating Spatial Visualization via Mathematical Transformations
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2026)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2026)
Brain Hematoma Marker Recognition Using Multitask Learning: SwinTransformer and Swin-Unet
von: Hirata, Kodai, et al.
Veröffentlicht: (2025)
von: Hirata, Kodai, et al.
Veröffentlicht: (2025)
A Transformer Based Handwriting Recognition System Jointly Using Online and Offline Features
von: Lodh, Ayush, et al.
Veröffentlicht: (2025)
von: Lodh, Ayush, et al.
Veröffentlicht: (2025)
Tab2Visual: Overcoming Limited Data in Tabular Data Classification Using Deep Learning with Visual Representations
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2025)
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
von: Vaish, Puru, et al.
Veröffentlicht: (2024) -
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
von: Berasi, Davide, et al.
Veröffentlicht: (2025) -
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
von: Cheng, Jintao, et al.
Veröffentlicht: (2025) -
Event-Driven Neuromorphic Vision Enables Energy-Efficient Visual Place Recognition
von: Keime, Geoffroy, et al.
Veröffentlicht: (2026) -
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
von: Ou, Yiwei, et al.
Veröffentlicht: (2025)