Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
Fuente:
arXiv
Saved in:
| Main Authors: | Metaxas, Ioannis Maniadis, Tzimiropoulos, Georgios, Patras, Ioannis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligned Unsupervised Pretraining of Object Detectors with Self-training
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
by: Feng, Chen, et al.
Published: (2021)
by: Feng, Chen, et al.
Published: (2021)
CLIPCleaner: Cleaning Noisy Labels with CLIP
by: Feng, Chen, et al.
Published: (2024)
by: Feng, Chen, et al.
Published: (2024)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
by: Bulat, Adrian, et al.
Published: (2026)
by: Bulat, Adrian, et al.
Published: (2026)
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
by: Das, Anurag, et al.
Published: (2026)
by: Das, Anurag, et al.
Published: (2026)
CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
by: Meng, Debin, et al.
Published: (2024)
by: Meng, Debin, et al.
Published: (2024)
LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
VladVA: Discriminative Fine-tuning of LVLMs
by: Ouali, Yassine, et al.
Published: (2024)
by: Ouali, Yassine, et al.
Published: (2024)
One-shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
Self-Supervised Facial Representation Learning with Facial Region Awareness
by: Gao, Zheng, et al.
Published: (2024)
by: Gao, Zheng, et al.
Published: (2024)
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
by: Xenos, Alexandros, et al.
Published: (2024)
by: Xenos, Alexandros, et al.
Published: (2024)
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
by: Krestenitis, Marios, et al.
Published: (2026)
by: Krestenitis, Marios, et al.
Published: (2026)
Prompting Visual-Language Models for Dynamic Facial Expression Recognition
by: Zhao, Zengqun, et al.
Published: (2023)
by: Zhao, Zengqun, et al.
Published: (2023)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
by: Meng, Debin, et al.
Published: (2025)
by: Meng, Debin, et al.
Published: (2025)
Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis
by: Feng, Chen, et al.
Published: (2026)
by: Feng, Chen, et al.
Published: (2026)
FOAA: Flattened Outer Arithmetic Attention For Multimodal Tumor Classification
by: Alwazzan, Omnia, et al.
Published: (2024)
by: Alwazzan, Omnia, et al.
Published: (2024)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
by: Singh, Abhishek Kumar, et al.
Published: (2024)
by: Singh, Abhishek Kumar, et al.
Published: (2024)
Are CLIP features all you need for Universal Synthetic Image Origin Attribution?
by: Cioni, Dario, et al.
Published: (2024)
by: Cioni, Dario, et al.
Published: (2024)
Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer
by: Zhao, Zengqun, et al.
Published: (2024)
by: Zhao, Zengqun, et al.
Published: (2024)
Behaviour4All: in-the-wild Facial Behaviour Analysis Toolkit
by: Kollias, Dimitrios, et al.
Published: (2024)
by: Kollias, Dimitrios, et al.
Published: (2024)
Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
by: Cao, Yu, et al.
Published: (2025)
by: Cao, Yu, et al.
Published: (2025)
Deep Clustering Using the Soft Silhouette Score: Towards Compact and Well-Separated Clusters
by: Vardakas, Georgios, et al.
Published: (2024)
by: Vardakas, Georgios, et al.
Published: (2024)
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
by: Nikandrou, Malvina, et al.
Published: (2024)
by: Nikandrou, Malvina, et al.
Published: (2024)
Visual inspection for illicit items in X-ray images using Deep Learning
by: Mademlis, Ioannis, et al.
Published: (2023)
by: Mademlis, Ioannis, et al.
Published: (2023)
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024)
by: Goulas, Andreas, et al.
Published: (2024)
P-TAME: Explain Any Image Classifier with Trained Perturbations
by: Ntrougkas, Mariano V., et al.
Published: (2025)
by: Ntrougkas, Mariano V., et al.
Published: (2025)
Cluster Contrast for Unsupervised Visual Representation Learning
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD
by: Ntinou, Ioanna, et al.
Published: (2024)
by: Ntinou, Ioanna, et al.
Published: (2024)
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
by: Ntinou, Ioanna, et al.
Published: (2023)
by: Ntinou, Ioanna, et al.
Published: (2023)
MOAB: Multi-Modal Outer Arithmetic Block For Fusion Of Histopathological Images And Genetic Data For Brain Tumor Grading
by: Alwazzan, Omnia, et al.
Published: (2024)
by: Alwazzan, Omnia, et al.
Published: (2024)
Fwd2Bot: LVLM Visual Token Compression with Double Forward Bottleneck
by: Bulat, Adrian, et al.
Published: (2025)
by: Bulat, Adrian, et al.
Published: (2025)
FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models
by: Sahili, Zahraa Al, et al.
Published: (2024)
by: Sahili, Zahraa Al, et al.
Published: (2024)
Visual Place Recognition for Large-Scale UAV Applications
by: Papapetros, Ioannis Tsampikos, et al.
Published: (2025)
by: Papapetros, Ioannis Tsampikos, et al.
Published: (2025)
Knowledge Distillation Meets Open-Set Semi-Supervised Learning
by: Yang, Jing, et al.
Published: (2022)
by: Yang, Jing, et al.
Published: (2022)
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
by: Zhao, Zengqun, et al.
Published: (2026)
by: Zhao, Zengqun, et al.
Published: (2026)
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
by: Bulat, Adrian, et al.
Published: (2024)
by: Bulat, Adrian, et al.
Published: (2024)
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
by: Ouali, Yassine, et al.
Published: (2024)
by: Ouali, Yassine, et al.
Published: (2024)
Similar Items
-
Aligned Unsupervised Pretraining of Object Detectors with Self-training
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023) -
SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
by: Feng, Chen, et al.
Published: (2021) -
CLIPCleaner: Cleaning Noisy Labels with CLIP
by: Feng, Chen, et al.
Published: (2024) -
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
by: Bulat, Adrian, et al.
Published: (2026) -
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
by: Das, Anurag, et al.
Published: (2026)