Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Hahyeon, Kwak, Nojun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909016217616384
author Choi, Hahyeon
Kwak, Nojun
author_facet Choi, Hahyeon
Kwak, Nojun
contents We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and shows a consistent reverse U-shaped sparsity-performance trend, with peak performance at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03348
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
Choi, Hahyeon
Kwak, Nojun
Machine Learning
Artificial Intelligence
We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and shows a consistent reverse U-shaped sparsity-performance trend, with peak performance at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches.
title Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.03348