Majority Voting for Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Launer, Tim, Hübotter, Jonas, Bagatella, Marco, Hakimi, Ido, Krause, Andreas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913040750870528
author Launer, Tim
Hübotter, Jonas
Bagatella, Marco
Hakimi, Ido
Krause, Andreas
author_facet Launer, Tim
Hübotter, Jonas
Bagatella, Marco
Hakimi, Ido
Krause, Andreas
contents We investigate Functional Majority Voting (FMV), a method based on functional consensus for code generation with Large Language Models, which identifies a representative solution from multiple generations using their runtime execution signatures on test inputs. We find that FMV is an effective test-time inference strategy, substantially boosting performance on LiveCodeBench without a large compute overhead. Furthermore, we extend the utility of functional consensus and apply it as an aggregation strategy for label-free Test-Time Reinforcement Learning. We demonstrate that this increases pass@1 on holdout tasks, but find no evidence of self-improvement beyond the base model's performance ceiling.
format Preprint
id arxiv_https___arxiv_org_abs_2604_15618
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Majority Voting for Code Generation
Launer, Tim
Hübotter, Jonas
Bagatella, Marco
Hakimi, Ido
Krause, Andreas
Machine Learning
We investigate Functional Majority Voting (FMV), a method based on functional consensus for code generation with Large Language Models, which identifies a representative solution from multiple generations using their runtime execution signatures on test inputs. We find that FMV is an effective test-time inference strategy, substantially boosting performance on LiveCodeBench without a large compute overhead. Furthermore, we extend the utility of functional consensus and apply it as an aggregation strategy for label-free Test-Time Reinforcement Learning. We demonstrate that this increases pass@1 on holdout tasks, but find no evidence of self-improvement beyond the base model's performance ceiling.
title Majority Voting for Code Generation
topic Machine Learning
url https://arxiv.org/abs/2604.15618