Tiny Inference-Time Scaling with Latent Verifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bucciarelli, Davide, Turri, Evelyn, Baraldi, Lorenzo, Cornia, Marcella, Cucchiara, Rita
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912981242085376
author Bucciarelli, Davide
Turri, Evelyn
Baraldi, Lorenzo
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
author_facet Bucciarelli, Davide
Turri, Evelyn
Baraldi, Lorenzo
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
contents Inference-time scaling has emerged as an effective way to improve generative models at test time by using a verifier to score and select candidate outputs. A common choice is to employ Multimodal Large Language Models (MLLMs) as verifiers, which can improve performance but introduce substantial inference-time cost. Indeed, diffusion pipelines operate in an autoencoder latent space to reduce computation, yet MLLM verifiers still require decoding candidates to pixel space and re-encoding them into the visual embedding space, leading to redundant and costly operations. In this work, we propose Verifier on Hidden States (VHS), a verifier that operates directly on intermediate hidden representations of Diffusion Transformer (DiT) single-step generators. VHS analyzes generator features without decoding to pixel space, thereby reducing the per-candidate verification cost while improving or matching the performance of MLLM-based competitors. We show that, under tiny inference budgets with only a small number of candidates per prompt, VHS enables more efficient inference-time scaling reducing joint generation-and-verification time by 63.3%, compute FLOPs by 51% and VRAM usage by 14.5% with respect to a standard MLLM verifier, achieving a +2.7% improvement on GenEval at the same inference-time budget.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22492
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tiny Inference-Time Scaling with Latent Verifiers
Bucciarelli, Davide
Turri, Evelyn
Baraldi, Lorenzo
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Inference-time scaling has emerged as an effective way to improve generative models at test time by using a verifier to score and select candidate outputs. A common choice is to employ Multimodal Large Language Models (MLLMs) as verifiers, which can improve performance but introduce substantial inference-time cost. Indeed, diffusion pipelines operate in an autoencoder latent space to reduce computation, yet MLLM verifiers still require decoding candidates to pixel space and re-encoding them into the visual embedding space, leading to redundant and costly operations. In this work, we propose Verifier on Hidden States (VHS), a verifier that operates directly on intermediate hidden representations of Diffusion Transformer (DiT) single-step generators. VHS analyzes generator features without decoding to pixel space, thereby reducing the per-candidate verification cost while improving or matching the performance of MLLM-based competitors. We show that, under tiny inference budgets with only a small number of candidates per prompt, VHS enables more efficient inference-time scaling reducing joint generation-and-verification time by 63.3%, compute FLOPs by 51% and VRAM usage by 14.5% with respect to a standard MLLM verifier, achieving a +2.7% improvement on GenEval at the same inference-time budget.
title Tiny Inference-Time Scaling with Latent Verifiers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2603.22492