Multi-label Cluster Discrimination for Visual Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Xiang, Yang, Kaicheng, Dai, Xiangzi, Feng, Ziyong, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
Region-based Cluster Discrimination for Visual Representation Learning
von: Xie, Yin, et al.
Veröffentlicht: (2025)
von: Xie, Yin, et al.
Veröffentlicht: (2025)
High-Fidelity Facial Albedo Estimation via Texture Quantization
von: Ran, Zimin, et al.
Veröffentlicht: (2024)
von: Ran, Zimin, et al.
Veröffentlicht: (2024)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
RWKV-CLIP: A Robust Vision-Language Representation Learner
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
von: Xie, Yin, et al.
Veröffentlicht: (2025)
von: Xie, Yin, et al.
Veröffentlicht: (2025)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
von: Xie, Yin, et al.
Veröffentlicht: (2024)
von: Xie, Yin, et al.
Veröffentlicht: (2024)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
von: Cui, Siying, et al.
Veröffentlicht: (2024)
von: Cui, Siying, et al.
Veröffentlicht: (2024)
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
von: Zheng, Tianlu, et al.
Veröffentlicht: (2025)
von: Zheng, Tianlu, et al.
Veröffentlicht: (2025)
1st Place Solution to the 1st SkatingVerse Challenge
von: Sun, Tao, et al.
Veröffentlicht: (2024)
von: Sun, Tao, et al.
Veröffentlicht: (2024)
Decoupled Global-Local Alignment for Improving Compositional Understanding
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
ForCenNet: Foreground-Centric Network for Document Image Rectification
von: Cai, Peng, et al.
Veröffentlicht: (2025)
von: Cai, Peng, et al.
Veröffentlicht: (2025)
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
von: Chen, Zhichao, et al.
Veröffentlicht: (2026)
von: Chen, Zhichao, et al.
Veröffentlicht: (2026)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
Neural Clustering based Visual Representation Learning
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
von: Chen, Guikun, et al.
Veröffentlicht: (2024)
Learning Representations for Clustering via Partial Information Discrimination and Cross-Level Interaction
von: Zhang, Hai-Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Hai-Xin, et al.
Veröffentlicht: (2024)
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
von: Tang, Feilong, et al.
Veröffentlicht: (2026)
MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation
von: Ye, Xinyan, et al.
Veröffentlicht: (2026)
von: Ye, Xinyan, et al.
Veröffentlicht: (2026)
Domain Adaptation for Multi-label Image Classification: a Discriminator-free Approach
von: Singh, Inder Pal, et al.
Veröffentlicht: (2025)
von: Singh, Inder Pal, et al.
Veröffentlicht: (2025)
Self-Supervised Discriminative Feature Learning for Deep Multi-View Clustering
von: Xu, Jie, et al.
Veröffentlicht: (2021)
von: Xu, Jie, et al.
Veröffentlicht: (2021)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
von: Dai, Ming, et al.
Veröffentlicht: (2025)
von: Dai, Ming, et al.
Veröffentlicht: (2025)
On the Discriminability of Self-Supervised Representation Learning
von: Song, Zeen, et al.
Veröffentlicht: (2024)
von: Song, Zeen, et al.
Veröffentlicht: (2024)
Pedestrian Attribute Recognition as Label-balanced Multi-label Learning
von: Zhou, Yibo, et al.
Veröffentlicht: (2024)
von: Zhou, Yibo, et al.
Veröffentlicht: (2024)
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2024)
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2024)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2023)
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2023)
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
von: Toker, Aysim, et al.
Veröffentlicht: (2025)
von: Toker, Aysim, et al.
Veröffentlicht: (2025)
Learning Disentangled Representations for Generalized Multi-view Clustering
von: Zou, Xin, et al.
Veröffentlicht: (2026)
von: Zou, Xin, et al.
Veröffentlicht: (2026)
Multi-label Instance-level Generalised Visual Grounding in Agriculture
von: Haghighat, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Haghighat, Mohammadreza, et al.
Veröffentlicht: (2026)
Revisiting Multi-Task Visual Representation Learning
von: Di, Shangzhe, et al.
Veröffentlicht: (2026)
von: Di, Shangzhe, et al.
Veröffentlicht: (2026)
Cluster Contrast for Unsupervised Visual Representation Learning
von: Giakoumoglou, Nikolaos, et al.
Veröffentlicht: (2025)
von: Giakoumoglou, Nikolaos, et al.
Veröffentlicht: (2025)
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
von: An, Xiang, et al.
Veröffentlicht: (2025)
von: An, Xiang, et al.
Veröffentlicht: (2025)
Weakly-supervised Localization of Manipulated Image Regions Using Multi-resolution Learned Features
von: Wang, Ziyong, et al.
Veröffentlicht: (2025)
von: Wang, Ziyong, et al.
Veröffentlicht: (2025)
Semantic Positive Pairs for Enhancing Visual Representation Learning of Instance Discrimination Methods
von: Alkhalefi, Mohammad, et al.
Veröffentlicht: (2023)
von: Alkhalefi, Mohammad, et al.
Veröffentlicht: (2023)
WaveFace: Authentic Face Restoration with Efficient Frequency Recovery
von: Miao, Yunqi, et al.
Veröffentlicht: (2024)
von: Miao, Yunqi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024) -
Region-based Cluster Discrimination for Visual Representation Learning
von: Xie, Yin, et al.
Veröffentlicht: (2025) -
High-Fidelity Facial Albedo Estimation via Texture Quantization
von: Ran, Zimin, et al.
Veröffentlicht: (2024) -
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
von: Zhang, Qian, et al.
Veröffentlicht: (2024) -
RWKV-CLIP: A Robust Vision-Language Representation Learner
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)