Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
Fuente:
arXiv
Saved in:
| Main Authors: | Ramirez-Garcia, Valeria, de-Fitero-Dominguez, David, Garcia-Cabot, Antonio, Garcia-Lopez, Eva |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
by: Fuster-Pena, Marcos, et al.
Published: (2025)
by: Fuster-Pena, Marcos, et al.
Published: (2025)
Evaluating Large Language Models for automatic analysis of teacher simulations
by: de-Fitero-Dominguez, David, et al.
Published: (2024)
by: de-Fitero-Dominguez, David, et al.
Published: (2024)
Enhanced Automated Code Vulnerability Repair using Large Language Models
by: de-Fitero-Dominguez, David, et al.
Published: (2024)
by: de-Fitero-Dominguez, David, et al.
Published: (2024)
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
by: Liu, Ying, et al.
Published: (2025)
by: Liu, Ying, et al.
Published: (2025)
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
by: Lei, Zhikai, et al.
Published: (2024)
by: Lei, Zhikai, et al.
Published: (2024)
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2024)
by: Tran, Hieu, et al.
Published: (2024)
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
Do LLMs estimate uncertainty well in instruction-following?
by: Heo, Juyeon, et al.
Published: (2024)
by: Heo, Juyeon, et al.
Published: (2024)
Do LLMs "know" internally when they follow instructions?
by: Heo, Juyeon, et al.
Published: (2024)
by: Heo, Juyeon, et al.
Published: (2024)
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs
by: Haq, Saiful, et al.
Published: (2025)
by: Haq, Saiful, et al.
Published: (2025)
LLMs can hide text in other text of the same length
by: Norelli, Antonio, et al.
Published: (2025)
by: Norelli, Antonio, et al.
Published: (2025)
Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
by: de Curtò, J., et al.
Published: (2025)
by: de Curtò, J., et al.
Published: (2025)
Abusive text transformation using LLMs
by: Chandra, Rohitash, et al.
Published: (2025)
by: Chandra, Rohitash, et al.
Published: (2025)
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2024)
by: Li, Dawei, et al.
Published: (2024)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
by: Stephan, Andreas, et al.
Published: (2024)
by: Stephan, Andreas, et al.
Published: (2024)
LLM-as-a-qualitative-judge: automating error analysis in natural language generation
by: Chirkova, Nadezhda, et al.
Published: (2025)
by: Chirkova, Nadezhda, et al.
Published: (2025)
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
by: Qian, Shun, et al.
Published: (2026)
by: Qian, Shun, et al.
Published: (2026)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
A Machine Learning-based Approach for Solving Recurrence Relations and its use in Cost Analysis of Logic Programs
by: Rustenholz, Louis, et al.
Published: (2024)
by: Rustenholz, Louis, et al.
Published: (2024)
Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
by: Zhang, Xiaotian, et al.
Published: (2025)
by: Zhang, Xiaotian, et al.
Published: (2025)
Among Them: A game-based framework for assessing persuasion capabilities of LLMs
by: Idziejczak, Mateusz, et al.
Published: (2025)
by: Idziejczak, Mateusz, et al.
Published: (2025)
Humans can learn to detect AI-generated texts, or at least learn when they can't
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
TIC: Translate-Infer-Compile for accurate "text to plan" using LLMs and Logical Representations
by: Agarwal, Sudhir, et al.
Published: (2024)
by: Agarwal, Sudhir, et al.
Published: (2024)
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
by: Buscemi, Alessio, et al.
Published: (2025)
by: Buscemi, Alessio, et al.
Published: (2025)
A survey on cutting-edge relation extraction techniques based on language models
by: Diaz-Garcia, Jose A., et al.
Published: (2024)
by: Diaz-Garcia, Jose A., et al.
Published: (2024)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
by: Cao, Hongliu, et al.
Published: (2025)
by: Cao, Hongliu, et al.
Published: (2025)
ClinText-SP and RigoBERTa Clinical: a new set of open resources for Spanish Clinical NLP
by: Subies, Guillem García, et al.
Published: (2025)
by: Subies, Guillem García, et al.
Published: (2025)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
by: Pletenev, Sergey, et al.
Published: (2025)
by: Pletenev, Sergey, et al.
Published: (2025)
Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
by: Ilić, David, et al.
Published: (2023)
by: Ilić, David, et al.
Published: (2023)
URL: Universal Referential Knowledge Linking via Task-instructed Representation Compression
by: Li, Zhuoqun, et al.
Published: (2024)
by: Li, Zhuoqun, et al.
Published: (2024)
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
by: Han, Yifu, et al.
Published: (2025)
by: Han, Yifu, et al.
Published: (2025)
Zero-shot cross-lingual transfer in instruction tuning of large language models
by: Chirkova, Nadezhda, et al.
Published: (2024)
by: Chirkova, Nadezhda, et al.
Published: (2024)
A Chat About Boring Problems: Studying GPT-based text normalization
by: Zhang, Yang, et al.
Published: (2023)
by: Zhang, Yang, et al.
Published: (2023)
Causality extraction from medical text using Large Language Models (LLMs)
by: Gopalakrishnan, Seethalakshmi, et al.
Published: (2024)
by: Gopalakrishnan, Seethalakshmi, et al.
Published: (2024)
LLMs left, right, and center: Assessing GPT's capabilities to label political bias from web domains
by: Hernandes, Raphael, et al.
Published: (2024)
by: Hernandes, Raphael, et al.
Published: (2024)
Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring
by: Dussert, Jehanne
Published: (2026)
by: Dussert, Jehanne
Published: (2026)
WizardLM: Empowering large pre-trained language models to follow complex instructions
by: Xu, Can, et al.
Published: (2023)
by: Xu, Can, et al.
Published: (2023)
Similar Items
-
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
by: de-Fitero-Dominguez, David, et al.
Published: (2025) -
RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
by: Fuster-Pena, Marcos, et al.
Published: (2025) -
Evaluating Large Language Models for automatic analysis of teacher simulations
by: de-Fitero-Dominguez, David, et al.
Published: (2024) -
Enhanced Automated Code Vulnerability Repair using Large Language Models
by: de-Fitero-Dominguez, David, et al.
Published: (2024) -
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
by: Liu, Ying, et al.
Published: (2025)