Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Hao, Liu, Yu, Yan, Shuanglin, Shen, Fei, He, Shengfeng, Qin, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
von: Song, Jiale, et al.
Veröffentlicht: (2026)
von: Song, Jiale, et al.
Veröffentlicht: (2026)
Deep Reversible Consistency Learning for Cross-modal Retrieval
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
von: Wu, Yi, et al.
Veröffentlicht: (2025)
von: Wu, Yi, et al.
Veröffentlicht: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
von: Yan, Yichen, et al.
Veröffentlicht: (2024)
von: Yan, Yichen, et al.
Veröffentlicht: (2024)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
von: Yu, Hang, et al.
Veröffentlicht: (2025)
von: Yu, Hang, et al.
Veröffentlicht: (2025)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
von: Jiang, Haoyu, et al.
Veröffentlicht: (2024)
von: Jiang, Haoyu, et al.
Veröffentlicht: (2024)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
Text-Only Data Synthesis for Vision Language Model Training
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
von: Li, Deng, et al.
Veröffentlicht: (2024)
von: Li, Deng, et al.
Veröffentlicht: (2024)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
von: Hao, Jing, et al.
Veröffentlicht: (2025)
von: Hao, Jing, et al.
Veröffentlicht: (2025)
GSSF: Generalized Structural Sparse Function for Deep Cross-modal Metric Learning
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
How Far Are We from Generating Missing Modalities with Foundation Models?
von: Ke, Guanzhou, et al.
Veröffentlicht: (2025)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2025)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
von: Xie, Liping, et al.
Veröffentlicht: (2025)
von: Xie, Liping, et al.
Veröffentlicht: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
von: Ding, Peng, et al.
Veröffentlicht: (2024)
von: Ding, Peng, et al.
Veröffentlicht: (2024)
Cross-modal Counterfactual Explanations: Uncovering Decision Factors and Dataset Biases in Subjective Classification
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
Unveiling Encoder-Free Vision-Language Models
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
Selective Vision-Language Subspace Projection for Few-shot CLIP
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
ComAlign: Compositional Alignment in Vision-Language Models
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
von: Tang, Hao, et al.
Veröffentlicht: (2025) -
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
von: Tang, Hao, et al.
Veröffentlicht: (2024) -
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024) -
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
von: Liu, Yu, et al.
Veröffentlicht: (2025) -
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)