Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hongjun, Hu, Po, Han, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiLo: A Learning Framework for Generalized Category Discovery Robust to Domain Shifts
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
von: Wang, Hongjun, et al.
Veröffentlicht: (2024)
Learning Part Knowledge to Facilitate Category Understanding for Fine-Grained Generalized Category Discovery
von: Wang, Enguang, et al.
Veröffentlicht: (2025)
von: Wang, Enguang, et al.
Veröffentlicht: (2025)
CFM: Language-aligned Concept Foundation Model for Vision
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
Weighted Risk Invariance: Domain Generalization under Invariant Feature Shift
von: Wong, Gina, et al.
Veröffentlicht: (2024)
von: Wong, Gina, et al.
Veröffentlicht: (2024)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection
von: Delibasoglu, Ibrahim
Veröffentlicht: (2026)
von: Delibasoglu, Ibrahim
Veröffentlicht: (2026)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
von: Lian, Chenyu, et al.
Veröffentlicht: (2025)
von: Lian, Chenyu, et al.
Veröffentlicht: (2025)
Scaling Semantic Categories: Investigating the Impact on Vision Transformer Labeling Performance
von: Lamelas, Anthony, et al.
Veröffentlicht: (2025)
von: Lamelas, Anthony, et al.
Veröffentlicht: (2025)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
FastVLM: Efficient Vision Encoding for Vision Language Models
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
von: Ma, Huan, et al.
Veröffentlicht: (2024)
von: Ma, Huan, et al.
Veröffentlicht: (2024)
Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
von: Xu, Qinwu
Veröffentlicht: (2026)
von: Xu, Qinwu
Veröffentlicht: (2026)
Hierarchical Generalized Category Discovery for Brain Tumor Classification in Digital Pathology
von: Perkonigg, Matthias, et al.
Veröffentlicht: (2025)
von: Perkonigg, Matthias, et al.
Veröffentlicht: (2025)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
von: Rastegar, Sarah, et al.
Veröffentlicht: (2024)
von: Rastegar, Sarah, et al.
Veröffentlicht: (2024)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
von: Yu, Geng, et al.
Veröffentlicht: (2024)
von: Yu, Geng, et al.
Veröffentlicht: (2024)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
von: Li, Kevin, et al.
Veröffentlicht: (2025)
von: Li, Kevin, et al.
Veröffentlicht: (2025)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
A Shift in Perspective on Causality in Domain Generalization
von: Machlanski, Damian, et al.
Veröffentlicht: (2025)
von: Machlanski, Damian, et al.
Veröffentlicht: (2025)
DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression
von: Saeed, Numan, et al.
Veröffentlicht: (2026)
von: Saeed, Numan, et al.
Veröffentlicht: (2026)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
von: Wang, Enguang, et al.
Veröffentlicht: (2024)
von: Wang, Enguang, et al.
Veröffentlicht: (2024)
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
von: Chua, Jia Yun, et al.
Veröffentlicht: (2025)
von: Chua, Jia Yun, et al.
Veröffentlicht: (2025)
Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models
von: Zhou, Andy, et al.
Veröffentlicht: (2023)
von: Zhou, Andy, et al.
Veröffentlicht: (2023)
PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
von: Maqbool, Danyal, et al.
Veröffentlicht: (2025)
von: Maqbool, Danyal, et al.
Veröffentlicht: (2025)
Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
von: Hossain, Shamima
Veröffentlicht: (2025)
von: Hossain, Shamima
Veröffentlicht: (2025)
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
von: Tang, Kai, et al.
Veröffentlicht: (2025)
von: Tang, Kai, et al.
Veröffentlicht: (2025)
Multi-Modal Adapter for Vision-Language Models
von: Seputis, Dominykas, et al.
Veröffentlicht: (2024)
von: Seputis, Dominykas, et al.
Veröffentlicht: (2024)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
Adapting Vision-Language Models for Evaluating World Models
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2025)
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2025)
PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
von: Aniraj, Ananthu, et al.
Veröffentlicht: (2024)
von: Aniraj, Ananthu, et al.
Veröffentlicht: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
von: Alkin, Benedikt, et al.
Veröffentlicht: (2024)
von: Alkin, Benedikt, et al.
Veröffentlicht: (2024)
SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
von: Caselli, Lorenzo, et al.
Veröffentlicht: (2026)
von: Caselli, Lorenzo, et al.
Veröffentlicht: (2026)
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2026)
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2026)
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
von: Hu, Yuanze, et al.
Veröffentlicht: (2025)
von: Hu, Yuanze, et al.
Veröffentlicht: (2025)
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
von: Schlarmann, Christian, et al.
Veröffentlicht: (2024)
von: Schlarmann, Christian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HiLo: A Learning Framework for Generalized Category Discovery Robust to Domain Shifts
von: Wang, Hongjun, et al.
Veröffentlicht: (2024) -
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
von: Sui, Elaine, et al.
Veröffentlicht: (2024) -
SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning
von: Wang, Hongjun, et al.
Veröffentlicht: (2024) -
Learning Part Knowledge to Facilitate Category Understanding for Fine-Grained Generalized Category Discovery
von: Wang, Enguang, et al.
Veröffentlicht: (2025) -
CFM: Language-aligned Concept Foundation Model for Vision
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)