Rethinking Positive Pairs in Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jiantao, Atito, Sara, Feng, Zhenhua, Mo, Shentong, Kitler, Josef, Awais, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DailyMAE: Towards Pretraining Masked Autoencoders in One Day
by: Wu, Jiantao, et al.
Published: (2024)
by: Wu, Jiantao, et al.
Published: (2024)
Investigating Self-Supervised Methods for Label-Efficient Learning
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
Channel-Aware Probing for Multi-Channel Imaging
by: Marikkar, Umar, et al.
Published: (2026)
by: Marikkar, Umar, et al.
Published: (2026)
C3R: Channel Conditioned Cell Representations for unified evaluation in microscopy imaging
by: Marikkar, Umar, et al.
Published: (2025)
by: Marikkar, Umar, et al.
Published: (2025)
Pseudo Labelling for Enhanced Masked Autoencoders
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection
by: Dong, Wenhua, et al.
Published: (2024)
by: Dong, Wenhua, et al.
Published: (2024)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
by: Mo, Shentong
Published: (2024)
by: Mo, Shentong
Published: (2024)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024)
by: Nazarieh, Fatemeh, et al.
Published: (2024)
C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
by: Li, Rongchang, et al.
Published: (2024)
by: Li, Rongchang, et al.
Published: (2024)
Domain Adaptation Without the Compute Burden for Efficient Whole Slide Image Analysis
by: Marikkar, Umar, et al.
Published: (2026)
by: Marikkar, Umar, et al.
Published: (2026)
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
by: Deria, Ankan, et al.
Published: (2025)
by: Deria, Ankan, et al.
Published: (2025)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
by: Mo, Shentong
Published: (2024)
by: Mo, Shentong
Published: (2024)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
GMAIL: Generative Modality Alignment for generated Image Learning
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Audio-visual Generalized Zero-shot Learning the Easy Way
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion
by: Cheng, Chunyang, et al.
Published: (2025)
by: Cheng, Chunyang, et al.
Published: (2025)
DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images
by: Marikkar, Umar, et al.
Published: (2026)
by: Marikkar, Umar, et al.
Published: (2026)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
DiCoM -- Diverse Concept Modeling towards Enhancing Generalizability in Chest X-Ray Studies
by: Parida, Abhijeet, et al.
Published: (2024)
by: Parida, Abhijeet, et al.
Published: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Semantic Grouping Network for Audio Source Separation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
by: Atito, Sara, et al.
Published: (2022)
by: Atito, Sara, et al.
Published: (2022)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
by: Mo, Shentong, et al.
Published: (2025)
by: Mo, Shentong, et al.
Published: (2025)
A Large-scale Medical Visual Task Adaptation Benchmark
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Rethink Arbitrary Style Transfer with Transformer and Contrastive Learning
by: Zhang, Zhanjie, et al.
Published: (2024)
by: Zhang, Zhanjie, et al.
Published: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
by: Nazarieh, Fatemeh, et al.
Published: (2025)
by: Nazarieh, Fatemeh, et al.
Published: (2025)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
Semantic Positive Pairs for Enhancing Visual Representation Learning of Instance Discrimination Methods
by: Alkhalefi, Mohammad, et al.
Published: (2023)
by: Alkhalefi, Mohammad, et al.
Published: (2023)
Data Augmentation of Contrastive Learning is Estimating Positive-incentive Noise
by: Zhang, Hongyuan, et al.
Published: (2024)
by: Zhang, Hongyuan, et al.
Published: (2024)
DASViT: Differentiable Architecture Search for Vision Transformer
by: Wu, Pengjin, et al.
Published: (2025)
by: Wu, Pengjin, et al.
Published: (2025)
Rethinking Image Forgery Detection via Soft Contrastive Learning and Unsupervised Clustering
by: Wu, Haiwei, et al.
Published: (2023)
by: Wu, Haiwei, et al.
Published: (2023)
Similar Items
-
DailyMAE: Towards Pretraining Masked Autoencoders in One Day
by: Wu, Jiantao, et al.
Published: (2024) -
Investigating Self-Supervised Methods for Label-Efficient Learning
by: Nandam, Srinivasa Rao, et al.
Published: (2024) -
Channel-Aware Probing for Multi-Channel Imaging
by: Marikkar, Umar, et al.
Published: (2026) -
C3R: Channel Conditioned Cell Representations for unified evaluation in microscopy imaging
by: Marikkar, Umar, et al.
Published: (2025) -
Pseudo Labelling for Enhanced Masked Autoencoders
by: Nandam, Srinivasa Rao, et al.
Published: (2024)