PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shirian, Melika, Vadaei, Kianoosh, Majlessi, Kian, Ebrahimi, Audrina, Hemmat, Arshia, Adibi, Peyman, Karshenas, Hossein
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918214862110720
author Shirian, Melika
Vadaei, Kianoosh
Majlessi, Kian
Ebrahimi, Audrina
Hemmat, Arshia
Adibi, Peyman
Karshenas, Hossein
author_facet Shirian, Melika
Vadaei, Kianoosh
Majlessi, Kian
Ebrahimi, Audrina
Hemmat, Arshia
Adibi, Peyman
Karshenas, Hossein
contents We present PrismSSL, a Python library that unifies state-of-the-art self-supervised learning (SSL) methods across audio, vision, graphs, and cross-modal settings in a single, modular codebase. The goal of the demo is to show how researchers and practitioners can: (i) install, configure, and run pretext training with a few lines of code; (ii) reproduce compact benchmarks; and (iii) extend the framework with new modalities or methods through clean trainer and dataset abstractions. PrismSSL is packaged on PyPI, released under the MIT license, integrates tightly with HuggingFace Transformers, and provides quality-of-life features such as distributed training in PyTorch, Optuna-based hyperparameter search, LoRA fine-tuning for Transformer backbones, animated embedding visualizations for sanity checks, Weights & Biases logging, and colorful, structured terminal logs for improved usability and clarity. In addition, PrismSSL offers a graphical dashboard - built with Flask and standard web technologies - that enables users to configure and launch training pipelines with minimal coding. The artifact (code and data recipes) will be publicly available and reproducible.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
Shirian, Melika
Vadaei, Kianoosh
Majlessi, Kian
Ebrahimi, Audrina
Hemmat, Arshia
Adibi, Peyman
Karshenas, Hossein
Machine Learning
Multimedia
We present PrismSSL, a Python library that unifies state-of-the-art self-supervised learning (SSL) methods across audio, vision, graphs, and cross-modal settings in a single, modular codebase. The goal of the demo is to show how researchers and practitioners can: (i) install, configure, and run pretext training with a few lines of code; (ii) reproduce compact benchmarks; and (iii) extend the framework with new modalities or methods through clean trainer and dataset abstractions. PrismSSL is packaged on PyPI, released under the MIT license, integrates tightly with HuggingFace Transformers, and provides quality-of-life features such as distributed training in PyTorch, Optuna-based hyperparameter search, LoRA fine-tuning for Transformer backbones, animated embedding visualizations for sanity checks, Weights & Biases logging, and colorful, structured terminal logs for improved usability and clarity. In addition, PrismSSL offers a graphical dashboard - built with Flask and standard web technologies - that enables users to configure and launch training pipelines with minimal coding. The artifact (code and data recipes) will be publicly available and reproducible.
title PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
topic Machine Learning
Multimedia
url https://arxiv.org/abs/2511.17776