Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909016217616384 |
|---|---|
| author | Choi, Hahyeon Kwak, Nojun |
| author_facet | Choi, Hahyeon Kwak, Nojun |
| contents | We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and shows a consistent reverse U-shaped sparsity-performance trend, with peak performance at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_03348 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts Choi, Hahyeon Kwak, Nojun Machine Learning Artificial Intelligence We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and shows a consistent reverse U-shaped sparsity-performance trend, with peak performance at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches. |
| title | Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2605.03348 |