NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tarasov, Denis, Nikulin, Alexander, Zisman, Ilya, Klepach, Albina, Lyubaykin, Nikita, Polubarov, Andrei, Derevyagin, Alexander, Kurenkov, Vladislav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912645531041792
author Tarasov, Denis
Nikulin, Alexander
Zisman, Ilya
Klepach, Albina
Lyubaykin, Nikita
Polubarov, Andrei
Derevyagin, Alexander
Kurenkov, Vladislav
author_facet Tarasov, Denis
Nikulin, Alexander
Zisman, Ilya
Klepach, Albina
Lyubaykin, Nikita
Polubarov, Andrei
Derevyagin, Alexander
Kurenkov, Vladislav
contents Recent advances in Vision-Language-Action (VLA) models have established a two-component architecture, where a pre-trained Vision-Language Model (VLM) encodes visual observations and task descriptions, and an action decoder maps these representations to continuous actions. Diffusion models have been widely adopted as action decoders due to their ability to model complex, multimodal action distributions. However, they require multiple iterative denoising steps at inference time or downstream techniques to speed up sampling, limiting their practicality in real-world settings where high-frequency control is crucial. In this work, we present NinA (Normalizing Flows in Action), a fast and expressive alternative to diffusion-based decoders for VLAs. NinA replaces the diffusion action decoder with a Normalizing Flow (NF) that enables one-shot sampling through an invertible transformation, significantly reducing inference time. We integrate NinA into the FLOWER VLA architecture and fine-tune on the LIBERO benchmark. Our experiments show that NinA matches the performance of its diffusion-based counterpart under the same training regime, while achieving substantially faster inference. These results suggest that NinA offers a promising path toward efficient, high-frequency VLA control without compromising performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16845
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
Tarasov, Denis
Nikulin, Alexander
Zisman, Ilya
Klepach, Albina
Lyubaykin, Nikita
Polubarov, Andrei
Derevyagin, Alexander
Kurenkov, Vladislav
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advances in Vision-Language-Action (VLA) models have established a two-component architecture, where a pre-trained Vision-Language Model (VLM) encodes visual observations and task descriptions, and an action decoder maps these representations to continuous actions. Diffusion models have been widely adopted as action decoders due to their ability to model complex, multimodal action distributions. However, they require multiple iterative denoising steps at inference time or downstream techniques to speed up sampling, limiting their practicality in real-world settings where high-frequency control is crucial. In this work, we present NinA (Normalizing Flows in Action), a fast and expressive alternative to diffusion-based decoders for VLAs. NinA replaces the diffusion action decoder with a Normalizing Flow (NF) that enables one-shot sampling through an invertible transformation, significantly reducing inference time. We integrate NinA into the FLOWER VLA architecture and fine-tune on the LIBERO benchmark. Our experiments show that NinA matches the performance of its diffusion-based counterpart under the same training regime, while achieving substantially faster inference. These results suggest that NinA offers a promising path toward efficient, high-frequency VLA control without compromising performance.
title NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.16845