Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Zijia, Zang, Zelin, Wang, Yelin, Yang, Guozheng, yu, Kaicheng, Chen, Wanyu, Wang, Miaoyu, Li, Stan Z. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Web-Scale Multimodal Summarization using CLIP-Based Semantic Alignment
von: K, Mounvik, et al.
Veröffentlicht: (2026)
von: K, Mounvik, et al.
Veröffentlicht: (2026)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
von: Yang, Shengzhu, et al.
Veröffentlicht: (2025)
von: Yang, Shengzhu, et al.
Veröffentlicht: (2025)
DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data Augmentation
von: Zang, Zelin, et al.
Veröffentlicht: (2023)
von: Zang, Zelin, et al.
Veröffentlicht: (2023)
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
von: Song, Wei, et al.
Veröffentlicht: (2025)
von: Song, Wei, et al.
Veröffentlicht: (2025)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
von: Ali, Muhammad, et al.
Veröffentlicht: (2024)
von: Ali, Muhammad, et al.
Veröffentlicht: (2024)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
GenURL: A General Framework for Unsupervised Representation Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2021)
von: Li, Siyuan, et al.
Veröffentlicht: (2021)
AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis
von: Tang, Tao, et al.
Veröffentlicht: (2024)
von: Tang, Tao, et al.
Veröffentlicht: (2024)
AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
von: Gao, Yijie, et al.
Veröffentlicht: (2025)
von: Gao, Yijie, et al.
Veröffentlicht: (2025)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
von: Yang, Zhiwei, et al.
Veröffentlicht: (2026)
von: Yang, Zhiwei, et al.
Veröffentlicht: (2026)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
von: Wang, Jun, et al.
Veröffentlicht: (2023)
von: Wang, Jun, et al.
Veröffentlicht: (2023)
Must: Maximizing Latent Capacity of Spatial Transcriptomics Data
von: Zang, Zelin, et al.
Veröffentlicht: (2024)
von: Zang, Zelin, et al.
Veröffentlicht: (2024)
Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal Alignment
von: Ma, Xiang, et al.
Veröffentlicht: (2026)
von: Ma, Xiang, et al.
Veröffentlicht: (2026)
CSFMamba: Cross State Fusion Mamba Operator for Multimodal Remote Sensing Image Classification
von: Wang, Qingyu, et al.
Veröffentlicht: (2025)
von: Wang, Qingyu, et al.
Veröffentlicht: (2025)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
von: Song, Dan, et al.
Veröffentlicht: (2023)
von: Song, Dan, et al.
Veröffentlicht: (2023)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings
von: Huang, Yue, et al.
Veröffentlicht: (2024)
von: Huang, Yue, et al.
Veröffentlicht: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic Understanding
von: Liao, Guibiao, et al.
Veröffentlicht: (2024)
von: Liao, Guibiao, et al.
Veröffentlicht: (2024)
Variance-Aware Loss Scheduling for Multimodal Alignment in Low-Data Settings
von: Pillai, Sneh
Veröffentlicht: (2025)
von: Pillai, Sneh
Veröffentlicht: (2025)
EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation
von: Zhang, Zelin, et al.
Veröffentlicht: (2025)
von: Zhang, Zelin, et al.
Veröffentlicht: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
USTEP: Spatio-Temporal Predictive Learning under A Unified View
von: Tan, Cheng, et al.
Veröffentlicht: (2023)
von: Tan, Cheng, et al.
Veröffentlicht: (2023)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
von: Das, Anurag, et al.
Veröffentlicht: (2024)
von: Das, Anurag, et al.
Veröffentlicht: (2024)
Distribution Aligned Semantics Adaption for Lifelong Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
von: Bakker, Hua Chang, et al.
Veröffentlicht: (2025)
von: Bakker, Hua Chang, et al.
Veröffentlicht: (2025)
HG3-NeRF: Hierarchical Geometric, Semantic, and Photometric Guided Neural Radiance Fields for Sparse View Inputs
von: Gao, Zelin, et al.
Veröffentlicht: (2024)
von: Gao, Zelin, et al.
Veröffentlicht: (2024)
DANCE: Dual-View Distribution Alignment for Dataset Condensation
von: Zhang, Hansong, et al.
Veröffentlicht: (2024)
von: Zhang, Hansong, et al.
Veröffentlicht: (2024)
Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
von: Song, Yizhi, et al.
Veröffentlicht: (2024)
von: Song, Yizhi, et al.
Veröffentlicht: (2024)
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
von: Zhang, Jiahuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahuan, et al.
Veröffentlicht: (2025)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
von: Song, Yehun, et al.
Veröffentlicht: (2025)
von: Song, Yehun, et al.
Veröffentlicht: (2025)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Semantic Alignment for Multimodal Large Language Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
SkinCLIP-VL: Consistency-Aware Vision-Language Learning for Multimodal Skin Cancer Diagnosis
von: Lu, Zhixiang, et al.
Veröffentlicht: (2026)
von: Lu, Zhixiang, et al.
Veröffentlicht: (2026)
Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
von: Sun, Xin, et al.
Veröffentlicht: (2025)
von: Sun, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Web-Scale Multimodal Summarization using CLIP-Based Semantic Alignment
von: K, Mounvik, et al.
Veröffentlicht: (2026) -
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
von: Yang, Shengzhu, et al.
Veröffentlicht: (2025) -
DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data Augmentation
von: Zang, Zelin, et al.
Veröffentlicht: (2023) -
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025) -
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)