Unified modality separation: A vision-language framework for unsupervised domain adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinyao, Li, Jingjing, Du, Zhekai, Zhu, Lei, Shen, Heng Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalizing vision-language models to novel domains: A comprehensive survey
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
by: Li, Xinyao, et al.
Published: (2024)
by: Li, Xinyao, et al.
Published: (2024)
Combining inherent knowledge of vision-language models with unsupervised domain adaptation through strong-weak guidance
by: Westfechtel, Thomas, et al.
Published: (2023)
by: Westfechtel, Thomas, et al.
Published: (2023)
Generalizing Vision-Language Models with Dedicated Prompt Guidance
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
BTMuda: A Bi-level Multi-source unsupervised domain adaptation framework for breast cancer diagnosis
by: Yang, Yuxiang, et al.
Published: (2024)
by: Yang, Yuxiang, et al.
Published: (2024)
Agile Multi-Source-Free Domain Adaptation
by: Li, Xinyao, et al.
Published: (2024)
by: Li, Xinyao, et al.
Published: (2024)
SADA: Semantic adversarial unsupervised domain adaptation for Temporal Action Localization
by: Pujol-Perich, David, et al.
Published: (2023)
by: Pujol-Perich, David, et al.
Published: (2023)
A multi-modal vision-language model for generalizable annotation-free pathology localization
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
SPGen: Stochastic scanpath generation for paintings using unsupervised domain adaptation
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
Automating construction safety inspections using a multi-modal vision-language RAG framework
by: Wang, Chenxin, et al.
Published: (2025)
by: Wang, Chenxin, et al.
Published: (2025)
Can language-guided unsupervised adaptation improve medical image classification using unpaired images and texts?
by: Rahman, Umaima, et al.
Published: (2024)
by: Rahman, Umaima, et al.
Published: (2024)
OmniColor: A Unified Framework for Multi-modal Lineart Colorization
by: Zhang, Xulu, et al.
Published: (2026)
by: Zhang, Xulu, et al.
Published: (2026)
bi-modal textual prompt learning for vision-language models in remote sensing
by: Kashyap, Pankhi, et al.
Published: (2026)
by: Kashyap, Pankhi, et al.
Published: (2026)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
by: Guo, Xiaotong, et al.
Published: (2025)
by: Guo, Xiaotong, et al.
Published: (2025)
Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
by: Yu, Bo, et al.
Published: (2025)
by: Yu, Bo, et al.
Published: (2025)
Exploring selective image matching methods for zero-shot and few-sample unsupervised domain adaptation of urban canopy prediction
by: Francis, John, et al.
Published: (2024)
by: Francis, John, et al.
Published: (2024)
SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
by: Zhang, Yingying, et al.
Published: (2025)
by: Zhang, Yingying, et al.
Published: (2025)
In defense of the two-stage framework for open-set domain adaptive semantic segmentation
by: Ren, Wenqi, et al.
Published: (2026)
by: Ren, Wenqi, et al.
Published: (2026)
Adaptive deep learning framework for robust unsupervised underwater image enhancement
by: Saleh, Alzayat, et al.
Published: (2022)
by: Saleh, Alzayat, et al.
Published: (2022)
Initialization matters in few-shot adaptation of vision-language models for histopathological image classification
by: Meseguer, Pablo, et al.
Published: (2026)
by: Meseguer, Pablo, et al.
Published: (2026)
From Galaxy Zoo DECaLS to BASS/MzLS: detailed galaxy morphology classification with unsupervised domain adaption
by: Ye, Renhao, et al.
Published: (2024)
by: Ye, Renhao, et al.
Published: (2024)
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
by: Liu, Chenyv, et al.
Published: (2026)
by: Liu, Chenyv, et al.
Published: (2026)
Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
GAQAT: gradient-adaptive quantization-aware training for domain generalization
by: Jiang, Jiacheng, et al.
Published: (2024)
by: Jiang, Jiacheng, et al.
Published: (2024)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
by: Nassar, Ahmed, et al.
Published: (2025)
by: Nassar, Ahmed, et al.
Published: (2025)
AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
UNICA: A Unified Neural Framework for Controllable 3D Avatars
by: Zhu, Jiahe, et al.
Published: (2026)
by: Zhu, Jiahe, et al.
Published: (2026)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
by: Tong, Wenwen, et al.
Published: (2025)
by: Tong, Wenwen, et al.
Published: (2025)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Enhancing medical vision-language contrastive learning via inter-matching relation modelling
by: Li, Mingjian, et al.
Published: (2024)
by: Li, Mingjian, et al.
Published: (2024)
PAGen: Phase-guided Amplitude Generation for Domain-adaptive Object Detection
by: Du, Shuchen, et al.
Published: (2025)
by: Du, Shuchen, et al.
Published: (2025)
A biologically inspired separable learning vision model for real-time traffic object perception in Dark
by: Li, Hulin, et al.
Published: (2025)
by: Li, Hulin, et al.
Published: (2025)
Openfly: A comprehensive platform for aerial vision-language navigation
by: Gao, Yunpeng, et al.
Published: (2025)
by: Gao, Yunpeng, et al.
Published: (2025)
DAS3D: Dual-modality Anomaly Synthesis for 3D Anomaly Detection
by: Li, Kecen, et al.
Published: (2024)
by: Li, Kecen, et al.
Published: (2024)
Unifying Physically-Informed Weather Priors in A Single Model for Image Restoration Across Multiple Adverse Weather Conditions
by: Xu, Jiaqi, et al.
Published: (2026)
by: Xu, Jiaqi, et al.
Published: (2026)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning
by: Li, Shuangzhi, et al.
Published: (2025)
by: Li, Shuangzhi, et al.
Published: (2025)
Unsupervised Spike Depth Estimation via Cross-modality Cross-domain Knowledge Transfer
by: Liu, Jiaming, et al.
Published: (2022)
by: Liu, Jiaming, et al.
Published: (2022)
Similar Items
-
Generalizing vision-language models to novel domains: A comprehensive survey
by: Li, Xinyao, et al.
Published: (2025) -
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
by: Li, Xinyao, et al.
Published: (2024) -
Combining inherent knowledge of vision-language models with unsupervised domain adaptation through strong-weak guidance
by: Westfechtel, Thomas, et al.
Published: (2023) -
Generalizing Vision-Language Models with Dedicated Prompt Guidance
by: Li, Xinyao, et al.
Published: (2025) -
BTMuda: A Bi-level Multi-source unsupervised domain adaptation framework for breast cancer diagnosis
by: Yang, Yuxiang, et al.
Published: (2024)