Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Lai, Peng, Zheng, Jianjie, Cheng, Sijie, Chen, Yun, Li, Peng, Liu, Yang, Chen, Guanhua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
di: Lai, Peng, et al.
Pubblicazione: (2026)
di: Lai, Peng, et al.
Pubblicazione: (2026)
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
di: Ruan, Zhiwen, et al.
Pubblicazione: (2026)
di: Ruan, Zhiwen, et al.
Pubblicazione: (2026)
Evaluating Memory Capability in Continuous Lifelog Scenario
di: Zheng, Jianjie, et al.
Pubblicazione: (2026)
di: Zheng, Jianjie, et al.
Pubblicazione: (2026)
Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
di: Xiao, Zeguan, et al.
Pubblicazione: (2025)
di: Xiao, Zeguan, et al.
Pubblicazione: (2025)
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
di: Wang, Yanli, et al.
Pubblicazione: (2026)
di: Wang, Yanli, et al.
Pubblicazione: (2026)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
di: Chen, Luyu, et al.
Pubblicazione: (2025)
di: Chen, Luyu, et al.
Pubblicazione: (2025)
Representation-Guided Parameter-Efficient LLM Unlearning
di: Xiao, Zeguan, et al.
Pubblicazione: (2026)
di: Xiao, Zeguan, et al.
Pubblicazione: (2026)
Beyond Prompt Content: Enhancing LLM Performance via Content-Format Integrated Prompt Optimization
di: Liu, Yuanye, et al.
Pubblicazione: (2025)
di: Liu, Yuanye, et al.
Pubblicazione: (2025)
G2: Guided Generation for Enhanced Output Diversity in LLMs
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
di: Su, Junyou, et al.
Pubblicazione: (2026)
di: Su, Junyou, et al.
Pubblicazione: (2026)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
di: Xiao, Zeguan, et al.
Pubblicazione: (2026)
di: Xiao, Zeguan, et al.
Pubblicazione: (2026)
LayAlign: Enhancing Multilingual Reasoning in Large Language Models via Layer-Wise Adaptive Fusion and Alignment Strategy
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
di: Song, Mingyang, et al.
Pubblicazione: (2026)
di: Song, Mingyang, et al.
Pubblicazione: (2026)
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
di: Xiao, Zeguan, et al.
Pubblicazione: (2025)
di: Xiao, Zeguan, et al.
Pubblicazione: (2025)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
di: Bhan, Milan, et al.
Pubblicazione: (2025)
di: Bhan, Milan, et al.
Pubblicazione: (2025)
MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
di: Wang, Hanqing, et al.
Pubblicazione: (2024)
di: Wang, Hanqing, et al.
Pubblicazione: (2024)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)
di: Yang, Bo, et al.
Pubblicazione: (2026)
Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem
di: Xiao, Zeguan, et al.
Pubblicazione: (2026)
di: Xiao, Zeguan, et al.
Pubblicazione: (2026)
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2025)
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2025)
Speak It Out: Solving Symbol-Related Problems with Symbol-to-Language Conversion for Language Models
di: Wang, Yile, et al.
Pubblicazione: (2024)
di: Wang, Yile, et al.
Pubblicazione: (2024)
DEEM: Dynamic Experienced Expert Modeling for Stance Detection
di: Wang, Xiaolong, et al.
Pubblicazione: (2024)
di: Wang, Xiaolong, et al.
Pubblicazione: (2024)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
di: Li, Weiyue, et al.
Pubblicazione: (2026)
di: Li, Weiyue, et al.
Pubblicazione: (2026)
Beyond Words: Evaluating and Bridging Epistemic Divergence in User-Agent Interaction via Theory of Mind
di: Ruan, Minyuan, et al.
Pubblicazione: (2026)
di: Ruan, Minyuan, et al.
Pubblicazione: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
di: Cai, Yunna, et al.
Pubblicazione: (2025)
di: Cai, Yunna, et al.
Pubblicazione: (2025)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
di: Xiong, Boya, et al.
Pubblicazione: (2025)
di: Xiong, Boya, et al.
Pubblicazione: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
di: Ko, Ching-Yun, et al.
Pubblicazione: (2026)
di: Ko, Ching-Yun, et al.
Pubblicazione: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2024)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2024)
Context-Aware Sentiment Forecasting via LLM-based Multi-Perspective Role-Playing Agents
di: Man, Fanhang, et al.
Pubblicazione: (2025)
di: Man, Fanhang, et al.
Pubblicazione: (2025)
Beyond Entity Alignment: Towards Complete Knowledge Graph Alignment via Entity-Relation Synergy
di: Fang, Xiaohan, et al.
Pubblicazione: (2024)
di: Fang, Xiaohan, et al.
Pubblicazione: (2024)
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025)
di: Li, Qingquan, et al.
Pubblicazione: (2025)
Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning
di: Qin, Chengwei, et al.
Pubblicazione: (2023)
di: Qin, Chengwei, et al.
Pubblicazione: (2023)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
di: Wei, Hui, et al.
Pubblicazione: (2024)
di: Wei, Hui, et al.
Pubblicazione: (2024)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
di: Hou, Yutao, et al.
Pubblicazione: (2026)
di: Hou, Yutao, et al.
Pubblicazione: (2026)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
di: Wang, Leyao, et al.
Pubblicazione: (2026)
di: Wang, Leyao, et al.
Pubblicazione: (2026)
Humans or LLMs as the Judge? A Study on Judgement Biases
di: Chen, Guiming Hardy, et al.
Pubblicazione: (2024)
di: Chen, Guiming Hardy, et al.
Pubblicazione: (2024)
Anchored Supervised Fine-Tuning
di: Zhu, He, et al.
Pubblicazione: (2025)
di: Zhu, He, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
di: Lai, Peng, et al.
Pubblicazione: (2026) -
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
di: Ruan, Zhiwen, et al.
Pubblicazione: (2026) -
Evaluating Memory Capability in Continuous Lifelog Scenario
di: Zheng, Jianjie, et al.
Pubblicazione: (2026) -
Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025) -
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
di: Xiao, Zeguan, et al.
Pubblicazione: (2025)