CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918192380641280 |
|---|---|
| author | Hosseini, Peyman Bohdal, Ondrej Ceritli, Taha Castro, Ignacio Purver, Matthew Ozay, Mete Michieli, Umberto |
| author_facet | Hosseini, Peyman Bohdal, Ondrej Ceritli, Taha Castro, Ignacio Purver, Matthew Ozay, Mete Michieli, Umberto |
| contents | Test-time Reinforcement Learning (TTRL) has shown promise in adapting foundation models for complex tasks at test-time, resulting in large performance improvements. TTRL leverages an elegant two-phase sampling strategy: first, multi-sampling derives a pseudo-label via majority voting, while subsequent downsampling and reward-based fine-tuning encourages the model to explore and learn diverse valid solutions, with the pseudo-label modulating the reward signal. Meanwhile, in-context learning has been widely explored at inference time and demonstrated the ability to enhance model performance without weight updates. However, TTRL's two-phase sampling strategy under-utilizes contextual guidance, which can potentially improve pseudo-label accuracy in the initial exploitation phase while regulating exploration in the second. To address this, we propose context-guided TTRL (CG-TTRL), integrating context dynamically into both sampling phases and propose a method for efficient context selection for on-device applications. Our evaluations on mathematical and scientific QA benchmarks show CG-TTRL outperforms TTRL (e.g. additional 7% relative accuracy improvement over TTRL), while boosting efficiency by obtaining strong performance after only a few steps of test-time training (e.g. 8% relative improvement rather than 1% over TTRL after 3 steps). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_06430 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models Hosseini, Peyman Bohdal, Ondrej Ceritli, Taha Castro, Ignacio Purver, Matthew Ozay, Mete Michieli, Umberto Machine Learning Computation and Language I.2.7; I.5.4 Test-time Reinforcement Learning (TTRL) has shown promise in adapting foundation models for complex tasks at test-time, resulting in large performance improvements. TTRL leverages an elegant two-phase sampling strategy: first, multi-sampling derives a pseudo-label via majority voting, while subsequent downsampling and reward-based fine-tuning encourages the model to explore and learn diverse valid solutions, with the pseudo-label modulating the reward signal. Meanwhile, in-context learning has been widely explored at inference time and demonstrated the ability to enhance model performance without weight updates. However, TTRL's two-phase sampling strategy under-utilizes contextual guidance, which can potentially improve pseudo-label accuracy in the initial exploitation phase while regulating exploration in the second. To address this, we propose context-guided TTRL (CG-TTRL), integrating context dynamically into both sampling phases and propose a method for efficient context selection for on-device applications. Our evaluations on mathematical and scientific QA benchmarks show CG-TTRL outperforms TTRL (e.g. additional 7% relative accuracy improvement over TTRL), while boosting efficiency by obtaining strong performance after only a few steps of test-time training (e.g. 8% relative improvement rather than 1% over TTRL after 3 steps). |
| title | CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models |
| topic | Machine Learning Computation and Language I.2.7; I.5.4 |
| url | https://arxiv.org/abs/2511.06430 |