Salient Mask-Guided Vision Transformer for Fine-Grained Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Demidov, Dmitry, Sharif, Muhammad Hamza, Abdurahimov, Aliakbar, Cholakkal, Hisham, Khan, Fahad Shahbaz |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
by: Sheikh, Tooba Tehreem, et al.
Published: (2025)
by: Sheikh, Tooba Tehreem, et al.
Published: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
AIN: The Arabic INclusive Large Multimodal Model
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
CONDA: Condensed Deep Association Learning for Co-Salient Object Detection
by: Li, Long, et al.
Published: (2024)
by: Li, Long, et al.
Published: (2024)
AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning
by: Awais, Muhammad, et al.
Published: (2024)
by: Awais, Muhammad, et al.
Published: (2024)
Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
PARIS3D: Reasoning-based 3D Part Segmentation Using Large Multimodal Model
by: Kareem, Amrin, et al.
Published: (2024)
by: Kareem, Amrin, et al.
Published: (2024)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
MAviS: A Multimodal Conversational Assistant For Avian Species
by: Kryklyvets, Yevheniia, et al.
Published: (2026)
by: Kryklyvets, Yevheniia, et al.
Published: (2026)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
by: Deria, Ankan, et al.
Published: (2026)
by: Deria, Ankan, et al.
Published: (2026)
GLaMM: Pixel Grounding Large Multimodal Model
by: Rasheed, Hanoona, et al.
Published: (2023)
by: Rasheed, Hanoona, et al.
Published: (2023)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
by: Ishaq, Ayesha, et al.
Published: (2025)
by: Ishaq, Ayesha, et al.
Published: (2025)
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
by: Silva, Sathira, et al.
Published: (2025)
by: Silva, Sathira, et al.
Published: (2025)
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
by: Nawaz, Umair, et al.
Published: (2025)
by: Nawaz, Umair, et al.
Published: (2025)
Semi-supervised Open-World Object Detection
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
Saccadic Vision for Fine-Grained Visual Classification
by: Schmidt, Johann, et al.
Published: (2025)
by: Schmidt, Johann, et al.
Published: (2025)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
by: Thawakar, Omkar, et al.
Published: (2023)
by: Thawakar, Omkar, et al.
Published: (2023)
CDChat: A Large Multimodal Model for Remote Sensing Change Description
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
by: Bharadwaj, Rohit, et al.
Published: (2023)
by: Bharadwaj, Rohit, et al.
Published: (2023)
Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation
by: Liu, Yaohua, et al.
Published: (2025)
by: Liu, Yaohua, et al.
Published: (2025)
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
by: Paul, Dipanjyoti, et al.
Published: (2023)
by: Paul, Dipanjyoti, et al.
Published: (2023)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
by: Chen, Shiming, et al.
Published: (2024)
by: Chen, Shiming, et al.
Published: (2024)
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
by: Chowdhury, Arpita, et al.
Published: (2025)
by: Chowdhury, Arpita, et al.
Published: (2025)
Distilling Local Texture Features for Colorectal Tissue Classification in Low Data Regimes
by: Demidov, Dmitry, et al.
Published: (2024)
by: Demidov, Dmitry, et al.
Published: (2024)
MaskAdapt: Unsupervised Geometry-Aware Domain Adaptation Using Multimodal Contextual Learning and RGB-Depth Masking
by: Nadeem, Numair, et al.
Published: (2025)
by: Nadeem, Numair, et al.
Published: (2025)
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
by: Boudjoghra, Mohamed El Amine, et al.
Published: (2024)
by: Boudjoghra, Mohamed El Amine, et al.
Published: (2024)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification
by: Zhu, Armando, et al.
Published: (2024)
by: Zhu, Armando, et al.
Published: (2024)
Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship Classification
by: Lan, Long, et al.
Published: (2024)
by: Lan, Long, et al.
Published: (2024)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
Efficient Transformer for High Resolution Image Motion Deblurring
by: Akmaral, Amanturdieva, et al.
Published: (2025)
by: Akmaral, Amanturdieva, et al.
Published: (2025)
Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking
by: Ishaq, Ayesha, et al.
Published: (2024)
by: Ishaq, Ayesha, et al.
Published: (2024)
Learning Camouflaged Object Detection from Noisy Pseudo Label
by: Zhang, Jin, et al.
Published: (2024)
by: Zhang, Jin, et al.
Published: (2024)
Advancements in Crop Analysis through Deep Learning and Explainable AI
by: Khan, Hamza
Published: (2025)
by: Khan, Hamza
Published: (2025)
Fine-Grained ImageNet Classification in the Wild
by: Lymperaiou, Maria, et al.
Published: (2023)
by: Lymperaiou, Maria, et al.
Published: (2023)
FILA: Fine-Grained Vision Language Models
by: Zhu, Shiding, et al.
Published: (2024)
by: Zhu, Shiding, et al.
Published: (2024)
Semantically Informed Salient Regions Guided Radiology Report Generation
by: Hou, Zeyi, et al.
Published: (2025)
by: Hou, Zeyi, et al.
Published: (2025)
Similar Items
-
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
by: Sheikh, Tooba Tehreem, et al.
Published: (2025) -
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025) -
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
by: Kumar, Komal, et al.
Published: (2025) -
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
by: Noman, Mubashir, et al.
Published: (2024) -
AIN: The Arabic INclusive Large Multimodal Model
by: Heakl, Ahmed, et al.
Published: (2025)