Improve LLM-as-a-Judge Ability as a General Ability
Fuente:
arXiv
Guardado en:
| Autores principales: | Yu, Jiachen, Sun, Shaoning, Hu, Xiaohui, Yan, Jiaxu, Yu, Kaidong, Li, Xuelong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
por: Sun, Shaoning, et al.
Publicado: (2025)
por: Sun, Shaoning, et al.
Publicado: (2025)
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
por: Kang, Zhaolu, et al.
Publicado: (2025)
por: Kang, Zhaolu, et al.
Publicado: (2025)
Improving LLM Abilities in Idiomatic Translation
por: Donthi, Sundesh, et al.
Publicado: (2024)
por: Donthi, Sundesh, et al.
Publicado: (2024)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
por: Li, Kaiyuan, et al.
Publicado: (2026)
por: Li, Kaiyuan, et al.
Publicado: (2026)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
por: Chang, Jiayi, et al.
Publicado: (2025)
por: Chang, Jiayi, et al.
Publicado: (2025)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
por: Sun, Yu, et al.
Publicado: (2024)
por: Sun, Yu, et al.
Publicado: (2024)
Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents
por: Sun, Haoran, et al.
Publicado: (2025)
por: Sun, Haoran, et al.
Publicado: (2025)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
por: Liu, Shuliang, et al.
Publicado: (2025)
por: Liu, Shuliang, et al.
Publicado: (2025)
LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
por: Li, Lingyao, et al.
Publicado: (2026)
por: Li, Lingyao, et al.
Publicado: (2026)
TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time Series
por: Sun, Chenxi, et al.
Publicado: (2023)
por: Sun, Chenxi, et al.
Publicado: (2023)
SALMONN: Towards Generic Hearing Abilities for Large Language Models
por: Tang, Changli, et al.
Publicado: (2023)
por: Tang, Changli, et al.
Publicado: (2023)
Prompt-Level Reward Specifications for Open-Ended Post-Training
por: Weng, Zijun, et al.
Publicado: (2026)
por: Weng, Zijun, et al.
Publicado: (2026)
Understanding the Ability of LLMs to Handle Character-Level Perturbation
por: Zhuo, Anyuan, et al.
Publicado: (2025)
por: Zhuo, Anyuan, et al.
Publicado: (2025)
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
por: Yu, Le, et al.
Publicado: (2023)
por: Yu, Le, et al.
Publicado: (2023)
Scheming Ability in LLM-to-LLM Strategic Interactions
por: Pham, Thao
Publicado: (2025)
por: Pham, Thao
Publicado: (2025)
Diversity of Thought Improves Reasoning Abilities of LLMs
por: Naik, Ranjita, et al.
Publicado: (2023)
por: Naik, Ranjita, et al.
Publicado: (2023)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
por: Yang, Xuewei, et al.
Publicado: (2026)
por: Yang, Xuewei, et al.
Publicado: (2026)
Pre-trained Language Models Improve the Few-shot Prompt Ability of Decision Transformer
por: Yang, Yu, et al.
Publicado: (2024)
por: Yang, Yu, et al.
Publicado: (2024)
Predicting Emergent Abilities with Infinite Resolution Evaluation
por: Hu, Shengding, et al.
Publicado: (2023)
por: Hu, Shengding, et al.
Publicado: (2023)
Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation
por: Wu, Junhong, et al.
Publicado: (2024)
por: Wu, Junhong, et al.
Publicado: (2024)
Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns
por: DuSell, Brian, et al.
Publicado: (2023)
por: DuSell, Brian, et al.
Publicado: (2023)
Crystal: Illuminating LLM Abilities on Language and Code
por: Tao, Tianhua, et al.
Publicado: (2024)
por: Tao, Tianhua, et al.
Publicado: (2024)
Preference-Aware Memory Update for Long-Term LLM Agents
por: Sun, Haoran, et al.
Publicado: (2025)
por: Sun, Haoran, et al.
Publicado: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
por: Yang, Bo, et al.
Publicado: (2026)
por: Yang, Bo, et al.
Publicado: (2026)
Reasoning Does Not Necessarily Improve Role-Playing Ability
por: Feng, Xiachong, et al.
Publicado: (2025)
por: Feng, Xiachong, et al.
Publicado: (2025)
Exploring Language Model's Code Generation Ability with Auxiliary Functions
por: Lee, Seonghyeon, et al.
Publicado: (2024)
por: Lee, Seonghyeon, et al.
Publicado: (2024)
Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue
por: Gu, Jia-Chen, et al.
Publicado: (2024)
por: Gu, Jia-Chen, et al.
Publicado: (2024)
Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models
por: Chen, Zhipeng, et al.
Publicado: (2024)
por: Chen, Zhipeng, et al.
Publicado: (2024)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
por: Costarelli, Anthony, et al.
Publicado: (2024)
por: Costarelli, Anthony, et al.
Publicado: (2024)
Cost-Efficient Estimation of General Abilities Across Benchmarks
por: Krumdick, Michael, et al.
Publicado: (2026)
por: Krumdick, Michael, et al.
Publicado: (2026)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
por: Kuang, Jiayi, et al.
Publicado: (2025)
por: Kuang, Jiayi, et al.
Publicado: (2025)
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
por: Tang, Zhenheng, et al.
Publicado: (2025)
por: Tang, Zhenheng, et al.
Publicado: (2025)
PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs
por: Liu, An, et al.
Publicado: (2024)
por: Liu, An, et al.
Publicado: (2024)
Teaching Human Behavior Improves Content Understanding Abilities Of LLMs
por: Singh, Somesh, et al.
Publicado: (2024)
por: Singh, Somesh, et al.
Publicado: (2024)
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
por: Sun, Haoran, et al.
Publicado: (2024)
por: Sun, Haoran, et al.
Publicado: (2024)
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
por: Gurjar, Priya, et al.
Publicado: (2026)
por: Gurjar, Priya, et al.
Publicado: (2026)
Code Pretraining Improves Entity Tracking Abilities of Language Models
por: Kim, Najoung, et al.
Publicado: (2024)
por: Kim, Najoung, et al.
Publicado: (2024)
On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing
por: Liu, Yule, et al.
Publicado: (2024)
por: Liu, Yule, et al.
Publicado: (2024)
A Survey on Enhancing Causal Reasoning Ability of Large Language Models
por: Li, Xin, et al.
Publicado: (2025)
por: Li, Xin, et al.
Publicado: (2025)
Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience Learning
por: Sun, Zhoujian, et al.
Publicado: (2025)
por: Sun, Zhoujian, et al.
Publicado: (2025)
Ejemplares similares
-
S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
por: Sun, Shaoning, et al.
Publicado: (2025) -
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
por: Kang, Zhaolu, et al.
Publicado: (2025) -
Improving LLM Abilities in Idiomatic Translation
por: Donthi, Sundesh, et al.
Publicado: (2024) -
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
por: Li, Kaiyuan, et al.
Publicado: (2026) -
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
por: Chang, Jiayi, et al.
Publicado: (2025)