Do Large Language Models Judge Error Severity Like Humans?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Diege, Chen, Guanyi, Fan, Zhao, Cheng, Xiaorong, He, Tingting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Much Do LLMs Know About Chinese Zero Pronouns?
von: Li, Yifei, et al.
Veröffentlicht: (2026)
von: Li, Yifei, et al.
Veröffentlicht: (2026)
How Do People Quantify Naturally: Evidence from Mandarin Picture Description
von: Zhang, Yayun, et al.
Veröffentlicht: (2026)
von: Zhang, Yayun, et al.
Veröffentlicht: (2026)
CCNU at SemEval-2025 Task 3: Leveraging Internal and External Knowledge of Large Language Models for Multilingual Hallucination Annotation
von: Liu, Xu, et al.
Veröffentlicht: (2025)
von: Liu, Xu, et al.
Veröffentlicht: (2025)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
When Seekers Are Hard to Help: Evaluating Emotional Support Dialogue Systems in Worst-Case Interactions
von: Yang, Jiajie, et al.
Veröffentlicht: (2026)
von: Yang, Jiajie, et al.
Veröffentlicht: (2026)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Large Language Models Are Human-Like Internally
von: Kuribayashi, Tatsuki, et al.
Veröffentlicht: (2025)
von: Kuribayashi, Tatsuki, et al.
Veröffentlicht: (2025)
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
Emotional Supporters often Use Multiple Strategies in a Single Turn
von: Bai, Xin, et al.
Veröffentlicht: (2025)
von: Bai, Xin, et al.
Veröffentlicht: (2025)
On the Robustness of Knowledge Editing for Detoxification
von: Dong, Ming, et al.
Veröffentlicht: (2026)
von: Dong, Ming, et al.
Veröffentlicht: (2026)
Can Large Language Models Express Uncertainty Like Human?
von: Tao, Linwei, et al.
Veröffentlicht: (2025)
von: Tao, Linwei, et al.
Veröffentlicht: (2025)
Rich Semantic Knowledge Enhanced Large Language Models for Few-shot Chinese Spell Checking
von: Dong, Ming, et al.
Veröffentlicht: (2024)
von: Dong, Ming, et al.
Veröffentlicht: (2024)
How Do Humans Write Code? Large Models Do It the Same Way Too
von: Li, Long, et al.
Veröffentlicht: (2024)
von: Li, Long, et al.
Veröffentlicht: (2024)
Do Large Language Models Solve ARC Visual Analogies Like People Do?
von: Opiełka, Gustaw, et al.
Veröffentlicht: (2024)
von: Opiełka, Gustaw, et al.
Veröffentlicht: (2024)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
Enhancing Human-Like Responses in Large Language Models
von: Çalık, Ethem Yağız, et al.
Veröffentlicht: (2025)
von: Çalık, Ethem Yağız, et al.
Veröffentlicht: (2025)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Does a Large Language Model Really Speak in Human-Like Language?
von: Park, Mose, et al.
Veröffentlicht: (2025)
von: Park, Mose, et al.
Veröffentlicht: (2025)
Do Influence Functions Work on Large Language Models?
von: Li, Zhe, et al.
Veröffentlicht: (2024)
von: Li, Zhe, et al.
Veröffentlicht: (2024)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
von: Zhao, Yuwei, et al.
Veröffentlicht: (2024)
von: Zhao, Yuwei, et al.
Veröffentlicht: (2024)
FedJudge: Federated Legal Large Language Model
von: Yue, Linan, et al.
Veröffentlicht: (2023)
von: Yue, Linan, et al.
Veröffentlicht: (2023)
Distortions in Judged Spatial Relations in Large Language Models
von: Fulman, Nir, et al.
Veröffentlicht: (2024)
von: Fulman, Nir, et al.
Veröffentlicht: (2024)
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
Do as We Do, Not as You Think: the Conformity of Large Language Models
von: Weng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Weng, Zhiyuan, et al.
Veröffentlicht: (2025)
Correct Like Humans: Progressive Learning Framework for Chinese Text Error Correction
von: Li, Yinghui, et al.
Veröffentlicht: (2023)
von: Li, Yinghui, et al.
Veröffentlicht: (2023)
A Survey on Human Preference Learning for Large Language Models
von: Jiang, Ruili, et al.
Veröffentlicht: (2024)
von: Jiang, Ruili, et al.
Veröffentlicht: (2024)
Lost in Translation: Do LVLM Judges Generalize Across Languages?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
Prompting Large Language Models with Human Error Markings for Self-Correcting Machine Translation
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
Computational Modelling of Plurality and Definiteness in Chinese Noun Phrases
von: Liu, Yuqi, et al.
Veröffentlicht: (2024)
von: Liu, Yuqi, et al.
Veröffentlicht: (2024)
Enhancing Training Data Attribution for Large Language Models with Fitting Error Consideration
von: Wu, Kangxi, et al.
Veröffentlicht: (2024)
von: Wu, Kangxi, et al.
Veröffentlicht: (2024)
LLMs Do Not Grade Essays Like Humans
von: Mathew, Jerin George, et al.
Veröffentlicht: (2026)
von: Mathew, Jerin George, et al.
Veröffentlicht: (2026)
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
von: Dash, Saloni, et al.
Veröffentlicht: (2025)
von: Dash, Saloni, et al.
Veröffentlicht: (2025)
Large Language Models Often Say One Thing and Do Another
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
Do Large Language Models Truly Understand Cross-cultural Differences?
von: Guo, Shiwei, et al.
Veröffentlicht: (2025)
von: Guo, Shiwei, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
When Large Language Models are Reliable for Judging Empathic Communication
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
von: Karakaş, Sercan
Veröffentlicht: (2026)
von: Karakaş, Sercan
Veröffentlicht: (2026)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
DSCD: Large Language Model Detoxification with Self-Constrained Decoding
von: Dong, Ming, et al.
Veröffentlicht: (2025)
von: Dong, Ming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Much Do LLMs Know About Chinese Zero Pronouns?
von: Li, Yifei, et al.
Veröffentlicht: (2026) -
How Do People Quantify Naturally: Evidence from Mandarin Picture Description
von: Zhang, Yayun, et al.
Veröffentlicht: (2026) -
CCNU at SemEval-2025 Task 3: Leveraging Internal and External Knowledge of Large Language Models for Multilingual Hallucination Annotation
von: Liu, Xu, et al.
Veröffentlicht: (2025) -
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
von: Lu, Qingyu, et al.
Veröffentlicht: (2023) -
When Seekers Are Hard to Help: Evaluating Emotional Support Dialogue Systems in Worst-Case Interactions
von: Yang, Jiajie, et al.
Veröffentlicht: (2026)