Towards an Improved Understanding and Utilization of Maximum Manifold Capacity Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schaeffer, Rylan, Lecomte, Victor, Pai, Dhruv Bhandarkar, Carranza, Andres, Isik, Berivan, Unell, Alyssa, Khona, Mikail, Yerxa, Thomas, LeCun, Yann, Chung, SueYeon, Gromov, Andrey, Shwartz-Ziv, Ravid, Koyejo, Sanmi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910485643788288
author Schaeffer, Rylan
Lecomte, Victor
Pai, Dhruv Bhandarkar
Carranza, Andres
Isik, Berivan
Unell, Alyssa
Khona, Mikail
Yerxa, Thomas
LeCun, Yann
Chung, SueYeon
Gromov, Andrey
Shwartz-Ziv, Ravid
Koyejo, Sanmi
author_facet Schaeffer, Rylan
Lecomte, Victor
Pai, Dhruv Bhandarkar
Carranza, Andres
Isik, Berivan
Unell, Alyssa
Khona, Mikail
Yerxa, Thomas
LeCun, Yann
Chung, SueYeon
Gromov, Andrey
Shwartz-Ziv, Ravid
Koyejo, Sanmi
contents Maximum Manifold Capacity Representations (MMCR) is a recent multi-view self-supervised learning (MVSSL) method that matches or surpasses other leading MVSSL methods. MMCR is intriguing because it does not fit neatly into any of the commonplace MVSSL lineages, instead originating from a statistical mechanical perspective on the linear separability of data manifolds. In this paper, we seek to improve our understanding and our utilization of MMCR. To better understand MMCR, we leverage tools from high dimensional probability to demonstrate that MMCR incentivizes alignment and uniformity of learned embeddings. We then leverage tools from information theory to show that such embeddings maximize a well-known lower bound on mutual information between views, thereby connecting the geometric perspective of MMCR to the information-theoretic perspective commonly discussed in MVSSL. To better utilize MMCR, we mathematically predict and experimentally confirm non-monotonic changes in the pretraining loss akin to double descent but with respect to atypical hyperparameters. We also discover compute scaling laws that enable predicting the pretraining loss as a function of gradients steps, batch size, embedding dimension and number of views. We then show that MMCR, originally applied to image data, is performant on multimodal image-text data. By more deeply understanding the theoretical and empirical behavior of MMCR, our work reveals insights on improving MVSSL methods.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09366
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards an Improved Understanding and Utilization of Maximum Manifold Capacity Representations
Schaeffer, Rylan
Lecomte, Victor
Pai, Dhruv Bhandarkar
Carranza, Andres
Isik, Berivan
Unell, Alyssa
Khona, Mikail
Yerxa, Thomas
LeCun, Yann
Chung, SueYeon
Gromov, Andrey
Shwartz-Ziv, Ravid
Koyejo, Sanmi
Machine Learning
Computer Vision and Pattern Recognition
Neurons and Cognition
Maximum Manifold Capacity Representations (MMCR) is a recent multi-view self-supervised learning (MVSSL) method that matches or surpasses other leading MVSSL methods. MMCR is intriguing because it does not fit neatly into any of the commonplace MVSSL lineages, instead originating from a statistical mechanical perspective on the linear separability of data manifolds. In this paper, we seek to improve our understanding and our utilization of MMCR. To better understand MMCR, we leverage tools from high dimensional probability to demonstrate that MMCR incentivizes alignment and uniformity of learned embeddings. We then leverage tools from information theory to show that such embeddings maximize a well-known lower bound on mutual information between views, thereby connecting the geometric perspective of MMCR to the information-theoretic perspective commonly discussed in MVSSL. To better utilize MMCR, we mathematically predict and experimentally confirm non-monotonic changes in the pretraining loss akin to double descent but with respect to atypical hyperparameters. We also discover compute scaling laws that enable predicting the pretraining loss as a function of gradients steps, batch size, embedding dimension and number of views. We then show that MMCR, originally applied to image data, is performant on multimodal image-text data. By more deeply understanding the theoretical and empirical behavior of MMCR, our work reveals insights on improving MVSSL methods.
title Towards an Improved Understanding and Utilization of Maximum Manifold Capacity Representations
topic Machine Learning
Computer Vision and Pattern Recognition
Neurons and Cognition
url https://arxiv.org/abs/2406.09366