PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Shirian, Melika, Vadaei, Kianoosh, Majlessi, Kian, Ebrahimi, Audrina, Hemmat, Arshia, Adibi, Peyman, Karshenas, Hossein |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval
por: Hemmat, Arshia, et al.
Publicado: (2024)
por: Hemmat, Arshia, et al.
Publicado: (2024)
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
por: Yan, Feng, et al.
Publicado: (2024)
por: Yan, Feng, et al.
Publicado: (2024)
Self-Supervised Image Super-Resolution Quality Assessment based on Content-Free Multi-Model Oriented Representation Learning
por: Majlessi, Kian, et al.
Publicado: (2026)
por: Majlessi, Kian, et al.
Publicado: (2026)
Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency
por: Zarghani, Abolfazl, et al.
Publicado: (2025)
por: Zarghani, Abolfazl, et al.
Publicado: (2025)
FORML: A Riemannian Hessian-free Method for Meta-learning on Stiefel Manifolds
por: Tabealhojeh, Hadi, et al.
Publicado: (2024)
por: Tabealhojeh, Hadi, et al.
Publicado: (2024)
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
por: Zhao, Yu, et al.
Publicado: (2025)
por: Zhao, Yu, et al.
Publicado: (2025)
TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities
por: Zhuang, Yan, et al.
Publicado: (2025)
por: Zhuang, Yan, et al.
Publicado: (2025)
ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection
por: Yu, Zihao, et al.
Publicado: (2025)
por: Yu, Zihao, et al.
Publicado: (2025)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
por: Cao, Jiajun, et al.
Publicado: (2025)
por: Cao, Jiajun, et al.
Publicado: (2025)
Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis
por: Li, Xiang, et al.
Publicado: (2026)
por: Li, Xiang, et al.
Publicado: (2026)
Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning
por: Xiong, Zechang, et al.
Publicado: (2026)
por: Xiong, Zechang, et al.
Publicado: (2026)
Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion Recognition
por: Nguyen, Cam-Van Thi, et al.
Publicado: (2024)
por: Nguyen, Cam-Van Thi, et al.
Publicado: (2024)
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
por: Zhang, Meishan, et al.
Publicado: (2024)
por: Zhang, Meishan, et al.
Publicado: (2024)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
por: Li, Qingcao, et al.
Publicado: (2026)
por: Li, Qingcao, et al.
Publicado: (2026)
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
por: Xu, Yang, et al.
Publicado: (2024)
por: Xu, Yang, et al.
Publicado: (2024)
TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Multimodal Interaction Modeling via Self-Supervised Multi-Task Learning for Review Helpfulness Prediction
por: Gong, HongLin, et al.
Publicado: (2024)
por: Gong, HongLin, et al.
Publicado: (2024)
Exploring Modality Disruption in Multimodal Fake News Detection
por: Liu, Moyang, et al.
Publicado: (2025)
por: Liu, Moyang, et al.
Publicado: (2025)
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
por: Tang, Anni, et al.
Publicado: (2022)
por: Tang, Anni, et al.
Publicado: (2022)
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
por: Li, Shuyu, et al.
Publicado: (2025)
por: Li, Shuyu, et al.
Publicado: (2025)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
por: An, Xiao, et al.
Publicado: (2026)
por: An, Xiao, et al.
Publicado: (2026)
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
por: Zou, Heqing, et al.
Publicado: (2024)
por: Zou, Heqing, et al.
Publicado: (2024)
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
por: Zhan, Hao, et al.
Publicado: (2026)
por: Zhan, Hao, et al.
Publicado: (2026)
Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
por: Ong, Rongqing Kenneth, et al.
Publicado: (2024)
por: Ong, Rongqing Kenneth, et al.
Publicado: (2024)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
por: Zhou, Qianrui, et al.
Publicado: (2023)
por: Zhou, Qianrui, et al.
Publicado: (2023)
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
por: Ma, Hongjian, et al.
Publicado: (2026)
por: Ma, Hongjian, et al.
Publicado: (2026)
Intelligent Carrier Allocation: A Cross-Modal Reasoning Framework for Adaptive Multimodal Steganography
por: Das, Abhirup, et al.
Publicado: (2025)
por: Das, Abhirup, et al.
Publicado: (2025)
Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft
por: Sun, Xin, et al.
Publicado: (2025)
por: Sun, Xin, et al.
Publicado: (2025)
Multi-level SSL Feature Gating for Audio Deepfake Detection
por: Tran, Hoan My, et al.
Publicado: (2025)
por: Tran, Hoan My, et al.
Publicado: (2025)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
por: Zhu, Xiaofei, et al.
Publicado: (2024)
por: Zhu, Xiaofei, et al.
Publicado: (2024)
A Multimodal Single-Branch Embedding Network for Recommendation in Cold-Start and Missing Modality Scenarios
por: Ganhör, Christian, et al.
Publicado: (2024)
por: Ganhör, Christian, et al.
Publicado: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
por: Zhang, Zhenxing, et al.
Publicado: (2024)
por: Zhang, Zhenxing, et al.
Publicado: (2024)
Multimodal Graph Neural Network for Recommendation with Dynamic De-redundancy and Modality-Guided Feature De-noisy
por: Mo, Feng, et al.
Publicado: (2024)
por: Mo, Feng, et al.
Publicado: (2024)
In-place Double Stimulus Methodology for Subjective Assessment of High Quality Images
por: Mohammadi, Shima, et al.
Publicado: (2025)
por: Mohammadi, Shima, et al.
Publicado: (2025)
M2ORT: Many-To-One Regression Transformer for Spatial Transcriptomics Prediction from Histopathology Images
por: Wang, Hongyi, et al.
Publicado: (2024)
por: Wang, Hongyi, et al.
Publicado: (2024)
EMID: An Emotional Aligned Dataset in Audio-Visual Modality
por: Zou, Jialing, et al.
Publicado: (2023)
por: Zou, Jialing, et al.
Publicado: (2023)
Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition
por: Zhuang, Yan, et al.
Publicado: (2026)
por: Zhuang, Yan, et al.
Publicado: (2026)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
por: Zhang, Hanlei, et al.
Publicado: (2024)
por: Zhang, Hanlei, et al.
Publicado: (2024)
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
por: Zheng, Xinhan, et al.
Publicado: (2025)
por: Zheng, Xinhan, et al.
Publicado: (2025)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
por: Zhou, Hengyang, et al.
Publicado: (2025)
por: Zhou, Hengyang, et al.
Publicado: (2025)
Ejemplares similares
-
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval
por: Hemmat, Arshia, et al.
Publicado: (2024) -
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
por: Yan, Feng, et al.
Publicado: (2024) -
Self-Supervised Image Super-Resolution Quality Assessment based on Content-Free Multi-Model Oriented Representation Learning
por: Majlessi, Kian, et al.
Publicado: (2026) -
Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency
por: Zarghani, Abolfazl, et al.
Publicado: (2025) -
FORML: A Riemannian Hessian-free Method for Meta-learning on Stiefel Manifolds
por: Tabealhojeh, Hadi, et al.
Publicado: (2024)