Image Recognition with Vision and Language Embeddings of VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Volkov, Illia, Kisel, Nikita, Janouskova, Klara, Matas, Jiri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Large Language Models as Image Classifiers
by: Kisel, Nikita, et al.
Published: (2026)
by: Kisel, Nikita, et al.
Published: (2026)
Flaws of ImageNet, Computer Vision's Favourite Dataset
by: Kisel, Nikita, et al.
Published: (2024)
by: Kisel, Nikita, et al.
Published: (2024)
Robust Context-Aware Object Recognition
by: Janouskova, Klara, et al.
Published: (2025)
by: Janouskova, Klara, et al.
Published: (2025)
Bringing the Context Back into Object Recognition, Robustly
by: Janouskova, Klara, et al.
Published: (2024)
by: Janouskova, Klara, et al.
Published: (2024)
Koo-Fu CLIP: Closed-Form Adaptation of Vision-Language Models via Fukunaga-Koontz Linear Discriminant Analysis
by: Suchanek, Matej, et al.
Published: (2026)
by: Suchanek, Matej, et al.
Published: (2026)
Single Image Test-Time Adaptation for Segmentation
by: Janouskova, Klara, et al.
Published: (2023)
by: Janouskova, Klara, et al.
Published: (2023)
FungiTastic: A multi-modal dataset and benchmark for image categorization
by: Picek, Lukas, et al.
Published: (2024)
by: Picek, Lukas, et al.
Published: (2024)
Breaking the Frame: Visual Place Recognition by Overlap Prediction
by: Wei, Tong, et al.
Published: (2024)
by: Wei, Tong, et al.
Published: (2024)
Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle
by: Purkrabek, Miroslav, et al.
Published: (2024)
by: Purkrabek, Miroslav, et al.
Published: (2024)
Video shutter angle estimation using optical flow and linear blur
by: Korcak, David, et al.
Published: (2023)
by: Korcak, David, et al.
Published: (2023)
Accurate Planar Tracking With Robust Re-Detection
by: Serych, Jonas, et al.
Published: (2026)
by: Serych, Jonas, et al.
Published: (2026)
Improving 2D Human Pose Estimation in Rare Camera Views with Synthetic Data
by: Purkrabek, Miroslav, et al.
Published: (2023)
by: Purkrabek, Miroslav, et al.
Published: (2023)
ProbPose: A Probabilistic Approach to 2D Human Pose Estimation
by: Purkrabek, Miroslav, et al.
Published: (2024)
by: Purkrabek, Miroslav, et al.
Published: (2024)
SAM2RL: Towards Reinforcement Learning Memory Control in Segment Anything Model 2
by: Adamyan, Alen, et al.
Published: (2025)
by: Adamyan, Alen, et al.
Published: (2025)
Human Pose-Constrained UV Map Estimation
by: Suchanek, Matej, et al.
Published: (2025)
by: Suchanek, Matej, et al.
Published: (2025)
Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection
by: Yermakov, Andrii, et al.
Published: (2025)
by: Yermakov, Andrii, et al.
Published: (2025)
SAM-pose2seg: Pose-Guided Human Instance Segmentation in Crowds
by: Kolomiiets, Constantin, et al.
Published: (2026)
by: Kolomiiets, Constantin, et al.
Published: (2026)
Dense Matchers for Dense Tracking
by: Jelínek, Tomáš, et al.
Published: (2024)
by: Jelínek, Tomáš, et al.
Published: (2024)
MFTIQ: Multi-Flow Tracker with Independent Matching Quality Estimation
by: Serych, Jonas, et al.
Published: (2024)
by: Serych, Jonas, et al.
Published: (2024)
Animal Identification with Independent Foreground and Background Modeling
by: Picek, Lukas, et al.
Published: (2024)
by: Picek, Lukas, et al.
Published: (2024)
PixOOD: Pixel-Level Out-of-Distribution Detection
by: Vojíř, Tomáš, et al.
Published: (2024)
by: Vojíř, Tomáš, et al.
Published: (2024)
BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
by: Purkrabek, Miroslav, et al.
Published: (2026)
by: Purkrabek, Miroslav, et al.
Published: (2026)
VPIT: Real-time Embedded Single Object 3D Tracking Using Voxel Pseudo Images
by: Oleksiienko, Illia, et al.
Published: (2022)
by: Oleksiienko, Illia, et al.
Published: (2022)
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
Deepfake Detection that Generalizes Across Benchmarks
by: Yermakov, Andrii, et al.
Published: (2025)
by: Yermakov, Andrii, et al.
Published: (2025)
HelixTrack: Event-Based Tracking and RPM Estimation of Propeller-like Objects
by: Spetlik, Radim, et al.
Published: (2026)
by: Spetlik, Radim, et al.
Published: (2026)
The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection
by: Yermakov, Andrii, et al.
Published: (2026)
by: Yermakov, Andrii, et al.
Published: (2026)
Global-Aware Edge Prioritization for Pose Graph Initialization
by: Wei, Tong, et al.
Published: (2026)
by: Wei, Tong, et al.
Published: (2026)
A New Dataset and a Distractor-Aware Architecture for Transparent Object Tracking
by: Lukezic, Alan, et al.
Published: (2024)
by: Lukezic, Alan, et al.
Published: (2024)
Three Things to Know about Deep Metric Learning
by: Patel, Yash, et al.
Published: (2024)
by: Patel, Yash, et al.
Published: (2024)
Going Beyond U-Net: Assessing Vision Transformers for Semantic Segmentation in Microscopy Image Analysis
by: Tsiporenko, Illia, et al.
Published: (2024)
by: Tsiporenko, Illia, et al.
Published: (2024)
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
Glo-VLMs: Leveraging Vision-Language Models for Fine-Grained Diseased Glomerulus Classification
by: Guo, Zhenhao, et al.
Published: (2025)
by: Guo, Zhenhao, et al.
Published: (2025)
What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models
by: Nie, Sen, et al.
Published: (2026)
by: Nie, Sen, et al.
Published: (2026)
Leveraging Vision-Language Embeddings for Zero-Shot Learning in Histopathology Images
by: Rahaman, Md Mamunur, et al.
Published: (2025)
by: Rahaman, Md Mamunur, et al.
Published: (2025)
VLEER: Vision and Language Embeddings for Explainable Whole Slide Image Representation
by: Nguyen, Anh Tien, et al.
Published: (2025)
by: Nguyen, Anh Tien, et al.
Published: (2025)
EE3P: Event-based Estimation of Periodic Phenomena Properties
by: Kolář, Jakub, et al.
Published: (2024)
by: Kolář, Jakub, et al.
Published: (2024)
EEPPR: Event-based Estimation of Periodic Phenomena Rate using Correlation in 3D
by: Kolář, Jakub, et al.
Published: (2024)
by: Kolář, Jakub, et al.
Published: (2024)
Point Cloud Color Constancy
by: Xing, Xiaoyan, et al.
Published: (2021)
by: Xing, Xiaoyan, et al.
Published: (2021)
WildFusion: Individual Animal Identification with Calibrated Similarity Fusion
by: Cermak, Vojtěch, et al.
Published: (2024)
by: Cermak, Vojtěch, et al.
Published: (2024)
Similar Items
-
Multimodal Large Language Models as Image Classifiers
by: Kisel, Nikita, et al.
Published: (2026) -
Flaws of ImageNet, Computer Vision's Favourite Dataset
by: Kisel, Nikita, et al.
Published: (2024) -
Robust Context-Aware Object Recognition
by: Janouskova, Klara, et al.
Published: (2025) -
Bringing the Context Back into Object Recognition, Robustly
by: Janouskova, Klara, et al.
Published: (2024) -
Koo-Fu CLIP: Closed-Form Adaptation of Vision-Language Models via Fukunaga-Koontz Linear Discriminant Analysis
by: Suchanek, Matej, et al.
Published: (2026)