Majority Voting for Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913040750870528 |
|---|---|
| author | Launer, Tim Hübotter, Jonas Bagatella, Marco Hakimi, Ido Krause, Andreas |
| author_facet | Launer, Tim Hübotter, Jonas Bagatella, Marco Hakimi, Ido Krause, Andreas |
| contents | We investigate Functional Majority Voting (FMV), a method based on functional consensus for code generation with Large Language Models, which identifies a representative solution from multiple generations using their runtime execution signatures on test inputs. We find that FMV is an effective test-time inference strategy, substantially boosting performance on LiveCodeBench without a large compute overhead. Furthermore, we extend the utility of functional consensus and apply it as an aggregation strategy for label-free Test-Time Reinforcement Learning. We demonstrate that this increases pass@1 on holdout tasks, but find no evidence of self-improvement beyond the base model's performance ceiling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_15618 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Majority Voting for Code Generation Launer, Tim Hübotter, Jonas Bagatella, Marco Hakimi, Ido Krause, Andreas Machine Learning We investigate Functional Majority Voting (FMV), a method based on functional consensus for code generation with Large Language Models, which identifies a representative solution from multiple generations using their runtime execution signatures on test inputs. We find that FMV is an effective test-time inference strategy, substantially boosting performance on LiveCodeBench without a large compute overhead. Furthermore, we extend the utility of functional consensus and apply it as an aggregation strategy for label-free Test-Time Reinforcement Learning. We demonstrate that this increases pass@1 on holdout tasks, but find no evidence of self-improvement beyond the base model's performance ceiling. |
| title | Majority Voting for Code Generation |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2604.15618 |