HEART: Emotionally-Driven Test-Time Scaling of Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910029619134464 |
|---|---|
| author | Pinto, Gabriela Goyal, Palash Parmar, Mihir Song, Yiwen Chakraborty, Souradip Wang, Zifeng Yoon, Jinsung Pfister, Tomas Palangi, Hamid |
| author_facet | Pinto, Gabriela Goyal, Palash Parmar, Mihir Song, Yiwen Chakraborty, Souradip Wang, Zifeng Yoon, Jinsung Pfister, Tomas Palangi, Hamid |
| contents | Test-time scaling has significantly improved how AI models solve problems, yet current methods often get stuck in repetitive, incorrect patterns of thought. We introduce HEART, a framework that uses emotional cues to guide the model's focus, much like how feelings contribute to human decision-making. By alternating between critical tones to sharpen error detection and encouraging tones to spark new ideas, HEART helps the model break out of dead-end reasoning and find the right solution. We evaluate HEART across seven high-difficulty benchmarks--including Humanity's Last Exam, GPQA Diamond, and LiveCodeBench--demonstrating robustness across diverse models. Results show that emotion facilitates deeper reasoning, yielding consistent accuracy gains over affect-sterile baselines. These findings suggest that the next frontier in machine reasoning lies in the strategic integration of affective regulation to guide logical synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_22876 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HEART: Emotionally-Driven Test-Time Scaling of Language Models Pinto, Gabriela Goyal, Palash Parmar, Mihir Song, Yiwen Chakraborty, Souradip Wang, Zifeng Yoon, Jinsung Pfister, Tomas Palangi, Hamid Computation and Language Machine Learning Test-time scaling has significantly improved how AI models solve problems, yet current methods often get stuck in repetitive, incorrect patterns of thought. We introduce HEART, a framework that uses emotional cues to guide the model's focus, much like how feelings contribute to human decision-making. By alternating between critical tones to sharpen error detection and encouraging tones to spark new ideas, HEART helps the model break out of dead-end reasoning and find the right solution. We evaluate HEART across seven high-difficulty benchmarks--including Humanity's Last Exam, GPQA Diamond, and LiveCodeBench--demonstrating robustness across diverse models. Results show that emotion facilitates deeper reasoning, yielding consistent accuracy gains over affect-sterile baselines. These findings suggest that the next frontier in machine reasoning lies in the strategic integration of affective regulation to guide logical synthesis. |
| title | HEART: Emotionally-Driven Test-Time Scaling of Language Models |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2509.22876 |