Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Maschan, Stepan, Qu, Haoxuan, Liu, Jun
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2601.03184
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914250328375296
author Maschan, Stepan
Qu, Haoxuan
Liu, Jun
author_facet Maschan, Stepan
Qu, Haoxuan
Liu, Jun
contents We present a theoretical analysis of decentralization of autoregressive generation. We define the Decentralized Discrete Flow Matching objective, by expressing probability generating velocity as a linear combination of expert flows. We also conduct experiments demonstrating the equivalence between decentralized and centralized training settings for multimodal language models across diverse set of benchmarks. Specifically, we compare two distinct paradigms: LLaVA and InternVL 2.5-1B, which uses a fixed CLIP vision encoder and performs full-parameter fine-tuning (ViT+MLP+LLM) during the instruction tuning stage.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03184
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Decentralized Autoregressive Generation
Maschan, Stepan
Qu, Haoxuan
Liu, Jun
Machine Learning
Artificial Intelligence
We present a theoretical analysis of decentralization of autoregressive generation. We define the Decentralized Discrete Flow Matching objective, by expressing probability generating velocity as a linear combination of expert flows. We also conduct experiments demonstrating the equivalence between decentralized and centralized training settings for multimodal language models across diverse set of benchmarks. Specifically, we compare two distinct paradigms: LLaVA and InternVL 2.5-1B, which uses a fixed CLIP vision encoder and performs full-parameter fine-tuning (ViT+MLP+LLM) during the instruction tuning stage.
title Decentralized Autoregressive Generation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.03184