HEART: Emotionally-Driven Test-Time Scaling of Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pinto, Gabriela, Goyal, Palash, Parmar, Mihir, Song, Yiwen, Chakraborty, Souradip, Wang, Zifeng, Yoon, Jinsung, Pfister, Tomas, Palangi, Hamid
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910029619134464
author Pinto, Gabriela
Goyal, Palash
Parmar, Mihir
Song, Yiwen
Chakraborty, Souradip
Wang, Zifeng
Yoon, Jinsung
Pfister, Tomas
Palangi, Hamid
author_facet Pinto, Gabriela
Goyal, Palash
Parmar, Mihir
Song, Yiwen
Chakraborty, Souradip
Wang, Zifeng
Yoon, Jinsung
Pfister, Tomas
Palangi, Hamid
contents Test-time scaling has significantly improved how AI models solve problems, yet current methods often get stuck in repetitive, incorrect patterns of thought. We introduce HEART, a framework that uses emotional cues to guide the model's focus, much like how feelings contribute to human decision-making. By alternating between critical tones to sharpen error detection and encouraging tones to spark new ideas, HEART helps the model break out of dead-end reasoning and find the right solution. We evaluate HEART across seven high-difficulty benchmarks--including Humanity's Last Exam, GPQA Diamond, and LiveCodeBench--demonstrating robustness across diverse models. Results show that emotion facilitates deeper reasoning, yielding consistent accuracy gains over affect-sterile baselines. These findings suggest that the next frontier in machine reasoning lies in the strategic integration of affective regulation to guide logical synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22876
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HEART: Emotionally-Driven Test-Time Scaling of Language Models
Pinto, Gabriela
Goyal, Palash
Parmar, Mihir
Song, Yiwen
Chakraborty, Souradip
Wang, Zifeng
Yoon, Jinsung
Pfister, Tomas
Palangi, Hamid
Computation and Language
Machine Learning
Test-time scaling has significantly improved how AI models solve problems, yet current methods often get stuck in repetitive, incorrect patterns of thought. We introduce HEART, a framework that uses emotional cues to guide the model's focus, much like how feelings contribute to human decision-making. By alternating between critical tones to sharpen error detection and encouraging tones to spark new ideas, HEART helps the model break out of dead-end reasoning and find the right solution. We evaluate HEART across seven high-difficulty benchmarks--including Humanity's Last Exam, GPQA Diamond, and LiveCodeBench--demonstrating robustness across diverse models. Results show that emotion facilitates deeper reasoning, yielding consistent accuracy gains over affect-sterile baselines. These findings suggest that the next frontier in machine reasoning lies in the strategic integration of affective regulation to guide logical synthesis.
title HEART: Emotionally-Driven Test-Time Scaling of Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.22876