Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Ashim, Srikumar, Vivek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916763144290304
author Gupta, Ashim
Srikumar, Vivek
author_facet Gupta, Ashim
Srikumar, Vivek
contents Inference-time scaling via repeated sampling has shown promise in reasoning tasks, but its effectiveness in multilingual generation remains underexplored. We evaluate this approach using perplexity- and reward-based verifiers on two multilingual benchmarks: the Aya Evaluation Suite and m-ArenaHard. Our results show consistent quality improvements, with gains exceeding 35% in some cases. While perplexity-based scoring is effective for open-ended prompts, only reward-based verifiers improve performance on tasks requiring reasoning (e.g., math, code). Our results demonstrate the broader utility of repeated sampling for multilingual text generation and underscore the importance of selecting right verifiers for the task.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21941
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
Gupta, Ashim
Srikumar, Vivek
Computation and Language
Inference-time scaling via repeated sampling has shown promise in reasoning tasks, but its effectiveness in multilingual generation remains underexplored. We evaluate this approach using perplexity- and reward-based verifiers on two multilingual benchmarks: the Aya Evaluation Suite and m-ArenaHard. Our results show consistent quality improvements, with gains exceeding 35% in some cases. While perplexity-based scoring is effective for open-ended prompts, only reward-based verifiers improve performance on tasks requiring reasoning (e.g., math, code). Our results demonstrate the broader utility of repeated sampling for multilingual text generation and underscore the importance of selecting right verifiers for the task.
title Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
topic Computation and Language
url https://arxiv.org/abs/2505.21941