LLM Benchmark-User Need Misalignment for Climate Change
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Oucheng, Xie, Lexing, Jiang, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health
by: Dey, Sumon Kanti, et al.
Published: (2025)
by: Dey, Sumon Kanti, et al.
Published: (2025)
Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
by: Liu, Zixuan, et al.
Published: (2025)
by: Liu, Zixuan, et al.
Published: (2025)
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
by: Zhang, Xiangxu, et al.
Published: (2026)
by: Zhang, Xiangxu, et al.
Published: (2026)
Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
by: Kurfalı, Murathan, et al.
Published: (2025)
by: Kurfalı, Murathan, et al.
Published: (2025)
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
by: Jiang, Xunyi, et al.
Published: (2025)
by: Jiang, Xunyi, et al.
Published: (2025)
BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge
by: Fu, Sihan, et al.
Published: (2026)
by: Fu, Sihan, et al.
Published: (2026)
An Investigation into Value Misalignment in LLM-Generated Texts for Cultural Heritage
by: Bu, Fan, et al.
Published: (2025)
by: Bu, Fan, et al.
Published: (2025)
Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings
by: Hong, Harbin, et al.
Published: (2025)
by: Hong, Harbin, et al.
Published: (2025)
CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing
by: Hu, Yucheng, et al.
Published: (2026)
by: Hu, Yucheng, et al.
Published: (2026)
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
by: Prama, Tabia Tanzin, et al.
Published: (2025)
by: Prama, Tabia Tanzin, et al.
Published: (2025)
Preemptive Detection and Correction of Misaligned Actions in LLM Agents
by: Fang, Haishuo, et al.
Published: (2024)
by: Fang, Haishuo, et al.
Published: (2024)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
User-LLM: Efficient LLM Contextualization with User Embeddings
by: Ning, Lin, et al.
Published: (2024)
by: Ning, Lin, et al.
Published: (2024)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
Taxonomy of User Needs and Actions
by: Shelby, Renee, et al.
Published: (2025)
by: Shelby, Renee, et al.
Published: (2025)
MoVa: Towards Generalizable Classification of Human Morals and Values
by: Chen, Ziyu, et al.
Published: (2025)
by: Chen, Ziyu, et al.
Published: (2025)
Retrieving Climate Change Disinformation by Narrative
by: Upravitelev, Max, et al.
Published: (2026)
by: Upravitelev, Max, et al.
Published: (2026)
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
How AI Forecasts AI Jobs: Benchmarking LLM Predictions of Labor Market Changes
by: Osborn, Sheri, et al.
Published: (2025)
by: Osborn, Sheri, et al.
Published: (2025)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
by: Zhang, Weizhi, et al.
Published: (2026)
by: Zhang, Weizhi, et al.
Published: (2026)
Strong Teacher Not Needed? On Distillation in LLM Pretraining
by: Lu, Taiming, et al.
Published: (2026)
by: Lu, Taiming, et al.
Published: (2026)
Climate Change from Large Language Models
by: Zhu, Hongyin, et al.
Published: (2023)
by: Zhu, Hongyin, et al.
Published: (2023)
Benchmark^2: Systematic Evaluation of LLM Benchmarks
by: Qian, Qi, et al.
Published: (2026)
by: Qian, Qi, et al.
Published: (2026)
Are Aligned Large Language Models Still Misaligned?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
Mitigating Temporal Misalignment by Discarding Outdated Facts
by: Zhang, Michael J. Q., et al.
Published: (2023)
by: Zhang, Michael J. Q., et al.
Published: (2023)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
by: Naik, Akshat, et al.
Published: (2025)
by: Naik, Akshat, et al.
Published: (2025)
MultiClimate: Multimodal Stance Detection on Climate Change Videos
by: Wang, Jiawen, et al.
Published: (2024)
by: Wang, Jiawen, et al.
Published: (2024)
Mitigating Misalignment Contagion by Steering with Implicit Traits
by: Chang, Maria, et al.
Published: (2026)
by: Chang, Maria, et al.
Published: (2026)
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
GMTRouter: Personalized LLM Router over Multi-turn User Interactions
by: Xie, Encheng, et al.
Published: (2025)
by: Xie, Encheng, et al.
Published: (2025)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
by: Wang, Yidong, et al.
Published: (2023)
by: Wang, Yidong, et al.
Published: (2023)
The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
by: Cao, Yixin, et al.
Published: (2025)
by: Cao, Yixin, et al.
Published: (2025)
MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
by: Lei, Yuzhen, et al.
Published: (2025)
by: Lei, Yuzhen, et al.
Published: (2025)
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation
by: Wu, Cheng-Kuang, et al.
Published: (2024)
by: Wu, Cheng-Kuang, et al.
Published: (2024)
Training on the Benchmark Is Not All You Need
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Similar Items
-
Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health
by: Dey, Sumon Kanti, et al.
Published: (2025) -
Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
by: Liu, Zixuan, et al.
Published: (2025) -
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
by: Zhang, Xiangxu, et al.
Published: (2026) -
Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
by: Kurfalı, Murathan, et al.
Published: (2025) -
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
by: Jiang, Xunyi, et al.
Published: (2025)