Salvato in:
| Autori principali: | Leyva-Vallina, María, Strisciuglio, Nicola, Petkov, Nicolai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2401.16304 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
di: Vaish, Puru, et al.
Pubblicazione: (2024)
di: Vaish, Puru, et al.
Pubblicazione: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
di: Berasi, Davide, et al.
Pubblicazione: (2025)
di: Berasi, Davide, et al.
Pubblicazione: (2025)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
di: Cheng, Jintao, et al.
Pubblicazione: (2025)
di: Cheng, Jintao, et al.
Pubblicazione: (2025)
Event-Driven Neuromorphic Vision Enables Energy-Efficient Visual Place Recognition
di: Keime, Geoffroy, et al.
Pubblicazione: (2026)
di: Keime, Geoffroy, et al.
Pubblicazione: (2026)
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
di: Ou, Yiwei, et al.
Pubblicazione: (2025)
di: Ou, Yiwei, et al.
Pubblicazione: (2025)
Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments
di: Shao, Xingyu, et al.
Pubblicazione: (2026)
di: Shao, Xingyu, et al.
Pubblicazione: (2026)
Revisit Anything: Visual Place Recognition via Image Segment Retrieval
di: Garg, Kartik, et al.
Pubblicazione: (2024)
di: Garg, Kartik, et al.
Pubblicazione: (2024)
Map-Relative Pose Regression for Visual Re-Localization
di: Chen, Shuai, et al.
Pubblicazione: (2024)
di: Chen, Shuai, et al.
Pubblicazione: (2024)
CART: Compositional Auto-Regressive Transformer for Image Generation
di: Roheda, Siddharth, et al.
Pubblicazione: (2024)
di: Roheda, Siddharth, et al.
Pubblicazione: (2024)
Contrastive Learning for Regression on Hyperspectral Data
di: Dhaini, Mohamad, et al.
Pubblicazione: (2024)
di: Dhaini, Mohamad, et al.
Pubblicazione: (2024)
Do ImageNet-trained models learn shortcuts? The impact of frequency shortcuts on generalization
di: Wang, Shunxin, et al.
Pubblicazione: (2025)
di: Wang, Shunxin, et al.
Pubblicazione: (2025)
GViT: Representing Images as Gaussians for Visual Recognition
di: Hernandez, Jefferson, et al.
Pubblicazione: (2025)
di: Hernandez, Jefferson, et al.
Pubblicazione: (2025)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
di: Du, Chaoqun, et al.
Pubblicazione: (2024)
di: Du, Chaoqun, et al.
Pubblicazione: (2024)
ConvMixFormer- A Resource-efficient Convolution Mixer for Transformer-based Dynamic Hand Gesture Recognition
di: Garg, Mallika, et al.
Pubblicazione: (2024)
di: Garg, Mallika, et al.
Pubblicazione: (2024)
Classification and regression of trajectories rendered as images via 2D Convolutional Neural Networks
di: Nicolai, Mariaclaudia, et al.
Pubblicazione: (2024)
di: Nicolai, Mariaclaudia, et al.
Pubblicazione: (2024)
EDTformer: An Efficient Decoder Transformer for Visual Place Recognition
di: Jin, Tong, et al.
Pubblicazione: (2024)
di: Jin, Tong, et al.
Pubblicazione: (2024)
Grounding Descriptions in Images informs Zero-Shot Visual Recognition
di: Halbe, Shaunak, et al.
Pubblicazione: (2024)
di: Halbe, Shaunak, et al.
Pubblicazione: (2024)
PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
di: Yang, Chenhongyi, et al.
Pubblicazione: (2024)
di: Yang, Chenhongyi, et al.
Pubblicazione: (2024)
Synthesizing Realistic Data for Table Recognition
di: Hou, Qiyu, et al.
Pubblicazione: (2024)
di: Hou, Qiyu, et al.
Pubblicazione: (2024)
TABLET: Table Structure Recognition using Encoder-only Transformers
di: Hou, Qiyu, et al.
Pubblicazione: (2025)
di: Hou, Qiyu, et al.
Pubblicazione: (2025)
Spectral-Spatial Contrastive Learning Framework for Regression on Hyperspectral Data
di: Dhaini, Mohamad, et al.
Pubblicazione: (2026)
di: Dhaini, Mohamad, et al.
Pubblicazione: (2026)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
di: Liang, Zichen, et al.
Pubblicazione: (2025)
di: Liang, Zichen, et al.
Pubblicazione: (2025)
Surely Large Multimodal Models (Don't) Excel in Visual Species Recognition?
di: Liu, Tian, et al.
Pubblicazione: (2025)
di: Liu, Tian, et al.
Pubblicazione: (2025)
Efficient Visual Transformer by Learnable Token Merging
di: Wang, Yancheng, et al.
Pubblicazione: (2024)
di: Wang, Yancheng, et al.
Pubblicazione: (2024)
Data-Driven Hierarchical Open Set Recognition
di: Hannum, Andrew, et al.
Pubblicazione: (2024)
di: Hannum, Andrew, et al.
Pubblicazione: (2024)
StrideNET: Swin Transformer for Terrain Recognition with Dynamic Roughness Extraction
di: Shelare, Maitreya, et al.
Pubblicazione: (2024)
di: Shelare, Maitreya, et al.
Pubblicazione: (2024)
Design and Analysis of Efficient Attention in Transformers for Social Group Activity Recognition
di: Tamura, Masato
Pubblicazione: (2024)
di: Tamura, Masato
Pubblicazione: (2024)
Minimalist Visual Inertial Odometry
di: Pasti, Francesco, et al.
Pubblicazione: (2026)
di: Pasti, Francesco, et al.
Pubblicazione: (2026)
Enhancing Tea Leaf Disease Recognition with Attention Mechanisms and Grad-CAM Visualization
di: Shikdar, Omar Faruq, et al.
Pubblicazione: (2025)
di: Shikdar, Omar Faruq, et al.
Pubblicazione: (2025)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
di: Zheng, Ying, et al.
Pubblicazione: (2025)
di: Zheng, Ying, et al.
Pubblicazione: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
di: Kuchibhotla, Hari Chandana, et al.
Pubblicazione: (2025)
di: Kuchibhotla, Hari Chandana, et al.
Pubblicazione: (2025)
Visualizing the loss landscape of Self-supervised Vision Transformer
di: Lee, Youngwan, et al.
Pubblicazione: (2024)
di: Lee, Youngwan, et al.
Pubblicazione: (2024)
HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimation
di: Zheng, Hongwei, et al.
Pubblicazione: (2025)
di: Zheng, Hongwei, et al.
Pubblicazione: (2025)
HATFormer: Historic Handwritten Arabic Text Recognition with Transformers
di: Chan, Adrian, et al.
Pubblicazione: (2024)
di: Chan, Adrian, et al.
Pubblicazione: (2024)
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
di: AlDahoul, Nouar, et al.
Pubblicazione: (2024)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2024)
Efficient License Plate Recognition in Videos Using Visual Rhythm and Accumulative Line Analysis
di: Ribeiro, Victor Nascimento, et al.
Pubblicazione: (2025)
di: Ribeiro, Victor Nascimento, et al.
Pubblicazione: (2025)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
di: Tang, Wei, et al.
Pubblicazione: (2025)
di: Tang, Wei, et al.
Pubblicazione: (2025)
TransformMix: Learning Transformation and Mixing Strategies from Data
di: Cheung, Tsz-Him, et al.
Pubblicazione: (2024)
di: Cheung, Tsz-Him, et al.
Pubblicazione: (2024)
Brain Hematoma Marker Recognition Using Multitask Learning: SwinTransformer and Swin-Unet
di: Hirata, Kodai, et al.
Pubblicazione: (2025)
di: Hirata, Kodai, et al.
Pubblicazione: (2025)
A Transformer Based Handwriting Recognition System Jointly Using Online and Offline Features
di: Lodh, Ayush, et al.
Pubblicazione: (2025)
di: Lodh, Ayush, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
di: Vaish, Puru, et al.
Pubblicazione: (2024) -
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
di: Berasi, Davide, et al.
Pubblicazione: (2025) -
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
di: Cheng, Jintao, et al.
Pubblicazione: (2025) -
Event-Driven Neuromorphic Vision Enables Energy-Efficient Visual Place Recognition
di: Keime, Geoffroy, et al.
Pubblicazione: (2026) -
MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
di: Ou, Yiwei, et al.
Pubblicazione: (2025)