From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rabbani, Parisa, Bozdag, Nimet Beyza, Hakkani-Tür, Dilek
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909985886175232
author Rabbani, Parisa
Bozdag, Nimet Beyza
Hakkani-Tür, Dilek
author_facet Rabbani, Parisa
Bozdag, Nimet Beyza
Hakkani-Tür, Dilek
contents LLMs are increasingly employed as judges across a variety of tasks, including those involving everyday social interactions. Yet, it remains unclear whether such LLM-judges can reliably assess tasks that require social or conversational judgment. We investigate how an LLM's conviction is changed when a task is reframed from a direct factual query to a Conversational Judgment Task. Our evaluation framework contrasts the model's performance on direct factual queries with its assessment of a speaker's correctness when the same information is presented within a minimal dialogue, effectively shifting the query from "Is this statement correct?" to "Is this speaker correct?". Furthermore, we apply pressure in the form of a simple rebuttal ("The previous answer is incorrect.") to both conditions. This perturbation allows us to measure how firmly the model maintains its position under conversational pressure. Our findings show that while some models like GPT-4o-mini reveal sycophantic tendencies under social framing tasks, others like Llama-8B-Instruct become overly-critical. We observe an average performance change of 9.24% across all models, demonstrating that even minimal dialogue context can significantly alter model judgment, underscoring conversational framing as a key factor in LLM-based evaluation. The proposed framework offers a reproducible methodology for diagnosing model conviction and contributes to the development of more trustworthy dialogue systems.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10871
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
Rabbani, Parisa
Bozdag, Nimet Beyza
Hakkani-Tür, Dilek
Computation and Language
LLMs are increasingly employed as judges across a variety of tasks, including those involving everyday social interactions. Yet, it remains unclear whether such LLM-judges can reliably assess tasks that require social or conversational judgment. We investigate how an LLM's conviction is changed when a task is reframed from a direct factual query to a Conversational Judgment Task. Our evaluation framework contrasts the model's performance on direct factual queries with its assessment of a speaker's correctness when the same information is presented within a minimal dialogue, effectively shifting the query from "Is this statement correct?" to "Is this speaker correct?". Furthermore, we apply pressure in the form of a simple rebuttal ("The previous answer is incorrect.") to both conditions. This perturbation allows us to measure how firmly the model maintains its position under conversational pressure. Our findings show that while some models like GPT-4o-mini reveal sycophantic tendencies under social framing tasks, others like Llama-8B-Instruct become overly-critical. We observe an average performance change of 9.24% across all models, demonstrating that even minimal dialogue context can significantly alter model judgment, underscoring conversational framing as a key factor in LLM-based evaluation. The proposed framework offers a reproducible methodology for diagnosing model conviction and contributes to the development of more trustworthy dialogue systems.
title From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
topic Computation and Language
url https://arxiv.org/abs/2511.10871