LLM Benchmark-User Need Misalignment for Climate Change

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Oucheng, Xie, Lexing, Jiang, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911547973959680
author Liu, Oucheng
Xie, Lexing
Jiang, Jing
author_facet Liu, Oucheng
Xie, Lexing
Jiang, Jing
contents Climate change is a major socio-scientific issue shapes public decision-making and policy discussions. As large language models (LLMs) increasingly serve as an interface for accessing climate knowledge, whether existing benchmarks reflect user needs is critical for evaluating LLM in real-world settings. We propose a Proactive Knowledge Behaviors Framework that captures the different human-human and human-AI knowledge seeking and provision behaviors. We further develop a Topic-Intent-Form taxonomy and apply it to analyze climate-related data representing different knowledge behaviors. Our results reveal a substantial mismatch between current benchmarks and real-world user needs, while knowledge interaction patterns between humans and LLMs closely resemble those in human-human interactions. These findings provide actionable guidance for benchmark design, RAG system development, and LLM training. Code is available at https://github.com/OuchengLiu/LLM-Misalign-Climate-Change.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26106
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM Benchmark-User Need Misalignment for Climate Change
Liu, Oucheng
Xie, Lexing
Jiang, Jing
Computation and Language
Climate change is a major socio-scientific issue shapes public decision-making and policy discussions. As large language models (LLMs) increasingly serve as an interface for accessing climate knowledge, whether existing benchmarks reflect user needs is critical for evaluating LLM in real-world settings. We propose a Proactive Knowledge Behaviors Framework that captures the different human-human and human-AI knowledge seeking and provision behaviors. We further develop a Topic-Intent-Form taxonomy and apply it to analyze climate-related data representing different knowledge behaviors. Our results reveal a substantial mismatch between current benchmarks and real-world user needs, while knowledge interaction patterns between humans and LLMs closely resemble those in human-human interactions. These findings provide actionable guidance for benchmark design, RAG system development, and LLM training. Code is available at https://github.com/OuchengLiu/LLM-Misalign-Climate-Change.
title LLM Benchmark-User Need Misalignment for Climate Change
topic Computation and Language
url https://arxiv.org/abs/2603.26106