Gespeichert in:
| Hauptverfasser: | Peng, Bo, Hu, Yuanwei, Liu, Bo, Chen, Ling, Lu, Jie, Fang, Zhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.09586 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models
von: Hu, Yuanwei, et al.
Veröffentlicht: (2026)
von: Hu, Yuanwei, et al.
Veröffentlicht: (2026)
On the Provable Importance of Gradients for Language-Assisted Image Clustering
von: Peng, Bo, et al.
Veröffentlicht: (2025)
von: Peng, Bo, et al.
Veröffentlicht: (2025)
Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models
von: Peng, Bo, et al.
Veröffentlicht: (2026)
von: Peng, Bo, et al.
Veröffentlicht: (2026)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
von: Liu, Shizhan, et al.
Veröffentlicht: (2025)
von: Liu, Shizhan, et al.
Veröffentlicht: (2025)
Delving into Out-of-Distribution Detection with Medical Vision-Language Models
von: Ju, Lie, et al.
Veröffentlicht: (2025)
von: Ju, Lie, et al.
Veröffentlicht: (2025)
CELLO: Causal Evaluation of Large Vision-Language Models
von: Chen, Meiqi, et al.
Veröffentlicht: (2024)
von: Chen, Meiqi, et al.
Veröffentlicht: (2024)
Negative Label Guided OOD Detection with Pretrained Vision-Language Models
von: Jiang, Xue, et al.
Veröffentlicht: (2024)
von: Jiang, Xue, et al.
Veröffentlicht: (2024)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
On the Learnability of Out-of-distribution Detection
von: Fang, Zhen, et al.
Veröffentlicht: (2024)
von: Fang, Zhen, et al.
Veröffentlicht: (2024)
DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
von: Zhao, Yuzhong, et al.
Veröffentlicht: (2024)
von: Zhao, Yuzhong, et al.
Veröffentlicht: (2024)
Multi-Token Enhancing for Vision Representation Learning
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2024)
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2024)
Delving Deep into Semantic Relation Distillation
von: Yan, Zhaoyi, et al.
Veröffentlicht: (2025)
von: Yan, Zhaoyi, et al.
Veröffentlicht: (2025)
MedSAM3: Delving into Segment Anything with Medical Concepts
von: Liu, Anglin, et al.
Veröffentlicht: (2025)
von: Liu, Anglin, et al.
Veröffentlicht: (2025)
FADE: Few-shot/zero-shot Anomaly Detection Engine using Large Vision-Language Model
von: Li, Yuanwei, et al.
Veröffentlicht: (2024)
von: Li, Yuanwei, et al.
Veröffentlicht: (2024)
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
von: Du, Hao, et al.
Veröffentlicht: (2025)
von: Du, Hao, et al.
Veröffentlicht: (2025)
HSCP: A Two-Stage Spectral Clustering Framework for Resource-Constrained UAV Identification
von: Wang, Maoyu, et al.
Veröffentlicht: (2025)
von: Wang, Maoyu, et al.
Veröffentlicht: (2025)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
von: Wang, Jiayu, et al.
Veröffentlicht: (2024)
von: Wang, Jiayu, et al.
Veröffentlicht: (2024)
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
von: Li, Ling, et al.
Veröffentlicht: (2026)
von: Li, Ling, et al.
Veröffentlicht: (2026)
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
von: Si, Dongchen, et al.
Veröffentlicht: (2025)
von: Si, Dongchen, et al.
Veröffentlicht: (2025)
Delving into Mapping Uncertainty for Mapless Trajectory Prediction
von: Zhang, Zongzheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zongzheng, et al.
Veröffentlicht: (2025)
CLIPSym: Delving into Symmetry Detection with CLIP
von: Yang, Tinghan, et al.
Veröffentlicht: (2025)
von: Yang, Tinghan, et al.
Veröffentlicht: (2025)
A Generative Framework for Self-Supervised Facial Representation Learning
von: He, Ruian, et al.
Veröffentlicht: (2023)
von: He, Ruian, et al.
Veröffentlicht: (2023)
Hyperspectral Image Classification via Efficient Global Spectral Supertoken Clustering
von: Liu, Peifu, et al.
Veröffentlicht: (2026)
von: Liu, Peifu, et al.
Veröffentlicht: (2026)
HGCLIP: Exploring Vision-Language Models with Graph Representations for Hierarchical Understanding
von: Xia, Peng, et al.
Veröffentlicht: (2023)
von: Xia, Peng, et al.
Veröffentlicht: (2023)
Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking
von: Ge, Jiawei, et al.
Veröffentlicht: (2023)
von: Ge, Jiawei, et al.
Veröffentlicht: (2023)
Delving into the Trajectory Long-tail Distribution for Muti-object Tracking
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks
von: Fu, Ronghao, et al.
Veröffentlicht: (2026)
von: Fu, Ronghao, et al.
Veröffentlicht: (2026)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
Bridging the Modality Gap in Roadside LiDAR: A Training-Free Vision-Language Model Framework for Vehicle Classification
von: Li, Yiqiao, et al.
Veröffentlicht: (2026)
von: Li, Yiqiao, et al.
Veröffentlicht: (2026)
Delving into Dark Regions for Robust Shadow Detection
von: Guan, Huankang, et al.
Veröffentlicht: (2024)
von: Guan, Huankang, et al.
Veröffentlicht: (2024)
CAS-IQA: Teaching Vision-Language Models for Synthetic Angiography Quality Assessment
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
Backdooring Vision-Language Models with Out-Of-Distribution Data
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
Dynamic Rank Adaptation for Vision-Language Models
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
TrojVLM: Backdoor Attack Against Vision Language Models
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model
von: Hu, Bing, et al.
Veröffentlicht: (2026)
von: Hu, Bing, et al.
Veröffentlicht: (2026)
SR$^{2}$-Net: A General Plug-and-Play Model for Spectral Refinement in Hyperspectral Image Super-Resolution
von: He, Ji-Xuan, et al.
Veröffentlicht: (2026)
von: He, Ji-Xuan, et al.
Veröffentlicht: (2026)
Balancing Complementarity and Consistency via Delayed Activation in Incomplete Multi-view Clustering
von: Li, Bo
Veröffentlicht: (2024)
von: Li, Bo
Veröffentlicht: (2024)
MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
von: Shi, Kuo, et al.
Veröffentlicht: (2025)
von: Shi, Kuo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models
von: Hu, Yuanwei, et al.
Veröffentlicht: (2026) -
On the Provable Importance of Gradients for Language-Assisted Image Clustering
von: Peng, Bo, et al.
Veröffentlicht: (2025) -
Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models
von: Peng, Bo, et al.
Veröffentlicht: (2026) -
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
von: Liu, Shizhan, et al.
Veröffentlicht: (2025) -
Delving into Out-of-Distribution Detection with Medical Vision-Language Models
von: Ju, Lie, et al.
Veröffentlicht: (2025)