From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signals
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Yeyong, Hu, Wenya, Wu, Xing, Qian, Quan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
by: Yu, Qingchen, et al.
Published: (2025)
by: Yu, Qingchen, et al.
Published: (2025)
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
by: Yu, Yeyong, et al.
Published: (2024)
by: Yu, Yeyong, et al.
Published: (2024)
AIMatDesign: Knowledge-Augmented Reinforcement Learning for Inverse Materials Design under Data Scarcity
by: Yu, Yeyong, et al.
Published: (2025)
by: Yu, Yeyong, et al.
Published: (2025)
From Small Data Modeling to Large Language Model Screening: A Dual‐Strategy Framework for Materials Intelligent Design
by: Yeyong Yu, et al.
Published: (2024)
by: Yeyong Yu, et al.
Published: (2024)
Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
by: de Langis, Karin, et al.
Published: (2025)
by: de Langis, Karin, et al.
Published: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
Beyond the Surface: Measuring Self-Preference in LLM Judgments
by: Chen, Zhi-Yuan, et al.
Published: (2025)
by: Chen, Zhi-Yuan, et al.
Published: (2025)
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
by: Khanmohammadi, Reza, et al.
Published: (2026)
by: Khanmohammadi, Reza, et al.
Published: (2026)
Good/Evil Reputation Judgment of Celebrities by LLMs via Retrieval Augmented Generation
by: Tsuchida, Rikuto, et al.
Published: (2025)
by: Tsuchida, Rikuto, et al.
Published: (2025)
Large Language Model-Enhanced Symbolic Reasoning for Knowledge Base Completion
by: He, Qiyuan, et al.
Published: (2025)
by: He, Qiyuan, et al.
Published: (2025)
How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
by: Ide, Yusuke, et al.
Published: (2024)
by: Ide, Yusuke, et al.
Published: (2024)
Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework
by: Ma, Xilai, et al.
Published: (2026)
by: Ma, Xilai, et al.
Published: (2026)
Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
by: Dentan, Jérémie, et al.
Published: (2025)
by: Dentan, Jérémie, et al.
Published: (2025)
Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models
by: Loginova, Olga, et al.
Published: (2024)
by: Loginova, Olga, et al.
Published: (2024)
From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations
by: Baskar, Sarvesh, et al.
Published: (2025)
by: Baskar, Sarvesh, et al.
Published: (2025)
Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
by: Lee, Wilson Y.
Published: (2026)
by: Lee, Wilson Y.
Published: (2026)
From Guessing to Placeholding: A Cost-Theoretic Framework for Uncertainty-Aware Code Completion
by: Zhu, Liang, et al.
Published: (2026)
by: Zhu, Liang, et al.
Published: (2026)
RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation
by: Yan, Shi-Qi, et al.
Published: (2025)
by: Yan, Shi-Qi, et al.
Published: (2025)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Permutative Preference Alignment from Listwise Ranking of Human Judgments
by: Zhao, Yang, et al.
Published: (2024)
by: Zhao, Yang, et al.
Published: (2024)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
Can Language Models Act as Knowledge Bases at Scale?
by: He, Qiyuan, et al.
Published: (2024)
by: He, Qiyuan, et al.
Published: (2024)
Reasoning over User Preferences: Knowledge Graph-Augmented LLMs for Explainable Conversational Recommendations
by: Qiu, Zhangchi, et al.
Published: (2024)
by: Qiu, Zhangchi, et al.
Published: (2024)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
by: Chen, Lu, et al.
Published: (2024)
by: Chen, Lu, et al.
Published: (2024)
Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning
by: Ni, Yijin, et al.
Published: (2025)
by: Ni, Yijin, et al.
Published: (2025)
Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
by: Lu, Yifan, et al.
Published: (2025)
by: Lu, Yifan, et al.
Published: (2025)
Evaluating and Aligning CodeLLMs on Human Preference
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
See or Guess: Counterfactually Regularized Image Captioning
by: Cao, Qian, et al.
Published: (2024)
by: Cao, Qian, et al.
Published: (2024)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
by: Tian, Changyuan, et al.
Published: (2025)
by: Tian, Changyuan, et al.
Published: (2025)
Benchmarking LLMs' Judgments with No Gold Standard
by: Xu, Shengwei, et al.
Published: (2024)
by: Xu, Shengwei, et al.
Published: (2024)
Enabling Discriminative Reasoning in LLMs for Legal Judgment Prediction
by: Deng, Chenlong, et al.
Published: (2024)
by: Deng, Chenlong, et al.
Published: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
RJE: A Retrieval-Judgment-Exploration Framework for Efficient Knowledge Graph Question Answering with LLMs
by: Lin, Can, et al.
Published: (2025)
by: Lin, Can, et al.
Published: (2025)
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
PACIFIC: Can LLMs Discern the Traits Influencing Your Preferences? Evaluating Personality-Driven Preference Alignment in LLMs
by: Zhao, Tianyu, et al.
Published: (2026)
by: Zhao, Tianyu, et al.
Published: (2026)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
Training Language Models to Generate Text with Citations via Fine-grained Rewards
by: Huang, Chengyu, et al.
Published: (2024)
by: Huang, Chengyu, et al.
Published: (2024)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
Similar Items
-
GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
by: Yu, Qingchen, et al.
Published: (2025) -
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
by: Yu, Yeyong, et al.
Published: (2024) -
AIMatDesign: Knowledge-Augmented Reinforcement Learning for Inverse Materials Design under Data Scarcity
by: Yu, Yeyong, et al.
Published: (2025) -
From Small Data Modeling to Large Language Model Screening: A Dual‐Strategy Framework for Materials Intelligent Design
by: Yeyong Yu, et al.
Published: (2024) -
Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
by: de Langis, Karin, et al.
Published: (2025)