XR-VLM: Cross-Relationship Modeling with Multi-part Prompts and Visual Features for Fine-Grained Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chuanming, Mao, Henming, Zhang, Huanhuan, Fu, Huiyuan, Ma, Huadong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Region-Aware Exposure Consistency Network for Mixed Exposure Correction
von: Liu, Jin, et al.
Veröffentlicht: (2024)
von: Liu, Jin, et al.
Veröffentlicht: (2024)
Mutual Distillation Learning For Person Re-Identification
von: Fu, Huiyuan, et al.
Veröffentlicht: (2024)
von: Fu, Huiyuan, et al.
Veröffentlicht: (2024)
Learning Exposure Correction in Dynamic Scenes
von: Liu, Jin, et al.
Veröffentlicht: (2024)
von: Liu, Jin, et al.
Veröffentlicht: (2024)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
von: Lin, Rui, et al.
Veröffentlicht: (2026)
von: Lin, Rui, et al.
Veröffentlicht: (2026)
Towards Efficient Object Re-Identification with A Novel Cloud-Edge Collaborative Framework
von: Wang, Chuanming, et al.
Veröffentlicht: (2024)
von: Wang, Chuanming, et al.
Veröffentlicht: (2024)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
von: Lin, Yuxin, et al.
Veröffentlicht: (2025)
von: Lin, Yuxin, et al.
Veröffentlicht: (2025)
On Learning Discriminative Features from Synthesized Data for Self-Supervised Fine-Grained Visual Recognition
von: Wang, Zihu, et al.
Veröffentlicht: (2024)
von: Wang, Zihu, et al.
Veröffentlicht: (2024)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
FG$^2$: Fine-Grained Cross-View Localization by Fine-Grained Feature Matching
von: Xia, Zimin, et al.
Veröffentlicht: (2025)
von: Xia, Zimin, et al.
Veröffentlicht: (2025)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
von: Chen, Tao, et al.
Veröffentlicht: (2025)
von: Chen, Tao, et al.
Veröffentlicht: (2025)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
von: Zhang, Qintong, et al.
Veröffentlicht: (2025)
von: Zhang, Qintong, et al.
Veröffentlicht: (2025)
LoDisc: Learning Global-Local Discriminative Features for Self-Supervised Fine-Grained Visual Recognition
von: Shi, Jialu, et al.
Veröffentlicht: (2024)
von: Shi, Jialu, et al.
Veröffentlicht: (2024)
M2Former: Multi-Scale Patch Selection for Fine-Grained Visual Recognition
von: Moon, Jiyong, et al.
Veröffentlicht: (2023)
von: Moon, Jiyong, et al.
Veröffentlicht: (2023)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
von: Ma, Yiming, et al.
Veröffentlicht: (2026)
von: Ma, Yiming, et al.
Veröffentlicht: (2026)
Fine-Grained VLM Fine-tuning via Latent Hierarchical Adapter Learning
von: Zhao, Yumiao, et al.
Veröffentlicht: (2025)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2025)
HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
von: Wu, Junhao, et al.
Veröffentlicht: (2025)
von: Wu, Junhao, et al.
Veröffentlicht: (2025)
Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual Categorization
von: Bi, Qi, et al.
Veröffentlicht: (2024)
von: Bi, Qi, et al.
Veröffentlicht: (2024)
EmbryoDiff: A Conditional Diffusion Framework with Multi-Focal Feature Fusion for Fine-Grained Embryo Developmental Stage Recognition
von: Sun, Yong, et al.
Veröffentlicht: (2025)
von: Sun, Yong, et al.
Veröffentlicht: (2025)
XR-VIO: High-precision Visual Inertial Odometry with Fast Initialization for XR Applications
von: Zhai, Shangjin, et al.
Veröffentlicht: (2025)
von: Zhai, Shangjin, et al.
Veröffentlicht: (2025)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
von: He, Wen-Jue, et al.
Veröffentlicht: (2025)
von: He, Wen-Jue, et al.
Veröffentlicht: (2025)
Micro-Expression Recognition via Fine-Grained Dynamic Perception
von: Shao, Zhiwen, et al.
Veröffentlicht: (2025)
von: Shao, Zhiwen, et al.
Veröffentlicht: (2025)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
Cross-Hierarchical Bidirectional Consistency Learning for Fine-Grained Visual Classification
von: Gao, Pengxiang, et al.
Veröffentlicht: (2025)
von: Gao, Pengxiang, et al.
Veröffentlicht: (2025)
Multi-Stage Contrastive Regression for Action Quality Assessment
von: An, Qi, et al.
Veröffentlicht: (2024)
von: An, Qi, et al.
Veröffentlicht: (2024)
Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action Recognition
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Enhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language Alignment
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
Feature Complementation Architecture for Visual Place Recognition
von: Wang, Weiwei, et al.
Veröffentlicht: (2025)
von: Wang, Weiwei, et al.
Veröffentlicht: (2025)
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
von: Shi, Liang, et al.
Veröffentlicht: (2026)
von: Shi, Liang, et al.
Veröffentlicht: (2026)
Cross-Block Fine-Grained Semantic Cascade for Skeleton-Based Sports Action Recognition
von: Liu, Zhendong, et al.
Veröffentlicht: (2024)
von: Liu, Zhendong, et al.
Veröffentlicht: (2024)
CVPT: Cross Visual Prompt Tuning
von: Huang, Lingyun, et al.
Veröffentlicht: (2024)
von: Huang, Lingyun, et al.
Veröffentlicht: (2024)
Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
von: Lian, Sheng, et al.
Veröffentlicht: (2025)
von: Lian, Sheng, et al.
Veröffentlicht: (2025)
A Large-Scale Remote Sensing Dataset and VLM-based Algorithm for Fine-Grained Road Hierarchy Classification
von: Han, Ting, et al.
Veröffentlicht: (2026)
von: Han, Ting, et al.
Veröffentlicht: (2026)
CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition
von: Zeeshan, Muhammad Osama, et al.
Veröffentlicht: (2026)
von: Zeeshan, Muhammad Osama, et al.
Veröffentlicht: (2026)
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
von: Zhang, Wei, et al.
Veröffentlicht: (2024)
von: Zhang, Wei, et al.
Veröffentlicht: (2024)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
von: He, Hulingxiao, et al.
Veröffentlicht: (2026)
von: He, Hulingxiao, et al.
Veröffentlicht: (2026)
Visual Prompting in LLMs for Enhancing Emotion Recognition
von: Zhang, Qixuan, et al.
Veröffentlicht: (2024)
von: Zhang, Qixuan, et al.
Veröffentlicht: (2024)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
von: Ma, Bo, et al.
Veröffentlicht: (2026)
von: Ma, Bo, et al.
Veröffentlicht: (2026)
Enhancing Fine-Grained Visual Recognition in the Low-Data Regime Through Feature Magnitude Regularization
von: Chapman, Avraham, et al.
Veröffentlicht: (2024)
von: Chapman, Avraham, et al.
Veröffentlicht: (2024)
Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
von: Pang, Cong, et al.
Veröffentlicht: (2025)
von: Pang, Cong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Region-Aware Exposure Consistency Network for Mixed Exposure Correction
von: Liu, Jin, et al.
Veröffentlicht: (2024) -
Mutual Distillation Learning For Person Re-Identification
von: Fu, Huiyuan, et al.
Veröffentlicht: (2024) -
Learning Exposure Correction in Dynamic Scenes
von: Liu, Jin, et al.
Veröffentlicht: (2024) -
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
von: Lin, Rui, et al.
Veröffentlicht: (2026) -
Towards Efficient Object Re-Identification with A Novel Cloud-Edge Collaborative Framework
von: Wang, Chuanming, et al.
Veröffentlicht: (2024)