Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2409.06754 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915013842698240 |
|---|---|
| author | Sun, Qingyun Guo, Zhen Team, PIN AI |
| author_facet | Sun, Qingyun Guo, Zhen Team, PIN AI |
| contents | We propose a scaling law hypothesis for multimodal models processing text, audio, images, and video within a shared token and embedding space. Our framework predicts model performance based on modality-specific compression and tokenization efficiency, extending established scaling laws from text-based decoder models to mixed-modality systems. We explore whether leveraging more training data in multiple modalities can reduce the size of the multimodal model, enabling efficient deployment on resource-constrained devices. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_06754 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Scaling Law Hypothesis for Multimodal Model Sun, Qingyun Guo, Zhen Team, PIN AI Machine Learning Artificial Intelligence We propose a scaling law hypothesis for multimodal models processing text, audio, images, and video within a shared token and embedding space. Our framework predicts model performance based on modality-specific compression and tokenization efficiency, extending established scaling laws from text-based decoder models to mixed-modality systems. We explore whether leveraging more training data in multiple modalities can reduce the size of the multimodal model, enabling efficient deployment on resource-constrained devices. |
| title | Scaling Law Hypothesis for Multimodal Model |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2409.06754 |