How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
Fuente:
arXiv
Saved in:
| Main Authors: | Doostmohammadi, Ehsan, Holmström, Oskar, Kuhlmann, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency
by: Doostmohammadi, Ehsan, et al.
Published: (2025)
by: Doostmohammadi, Ehsan, et al.
Published: (2025)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
by: Wang, Yidong, et al.
Published: (2023)
by: Wang, Yidong, et al.
Published: (2023)
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Fine-Tuning LLMs for Reliable Medical Question-Answering Services
by: Anaissi, Ali, et al.
Published: (2024)
by: Anaissi, Ali, et al.
Published: (2024)
Automatic Legal Writing Evaluation of LLMs
by: Pires, Ramon, et al.
Published: (2025)
by: Pires, Ramon, et al.
Published: (2025)
GemmAr: Enhancing LLMs Through Arabic Instruction-Tuning
by: Chouikhi, Hasna, et al.
Published: (2024)
by: Chouikhi, Hasna, et al.
Published: (2024)
How Reliable are LLMs for Reasoning on the Re-ranking task?
by: Islam, Nafis Tanveer, et al.
Published: (2025)
by: Islam, Nafis Tanveer, et al.
Published: (2025)
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
by: Chen, Yicheng, et al.
Published: (2025)
by: Chen, Yicheng, et al.
Published: (2025)
Towards Automatic Continual Learning: A Self-Adaptive Framework for Continual Instruction Tuning
by: Lin, Peiyi, et al.
Published: (2025)
by: Lin, Peiyi, et al.
Published: (2025)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
by: Zan, Changtong, et al.
Published: (2024)
by: Zan, Changtong, et al.
Published: (2024)
Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models
by: Kauf, Carina, et al.
Published: (2024)
by: Kauf, Carina, et al.
Published: (2024)
How LLMs Fail to Support Fact-Checking
by: Proma, Adiba Mahbub, et al.
Published: (2025)
by: Proma, Adiba Mahbub, et al.
Published: (2025)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
by: Arias-Duart, Anna, et al.
Published: (2025)
by: Arias-Duart, Anna, et al.
Published: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
by: Luo, Yifan, et al.
Published: (2024)
by: Luo, Yifan, et al.
Published: (2024)
A Comparative Analysis of Instruction Fine-Tuning LLMs for Financial Text Classification
by: Fatemi, Sorouralsadat, et al.
Published: (2024)
by: Fatemi, Sorouralsadat, et al.
Published: (2024)
Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
by: Verma, Pulkit, et al.
Published: (2025)
by: Verma, Pulkit, et al.
Published: (2025)
Instruction Tuning With Loss Over Instructions
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages
by: Mao, Zhuoyuan, et al.
Published: (2024)
by: Mao, Zhuoyuan, et al.
Published: (2024)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
by: Bajpai, Prasoon, et al.
Published: (2024)
by: Bajpai, Prasoon, et al.
Published: (2024)
Span-level Emotion-Cause-Category Triplet Extraction with Instruction Tuning LLMs and Data Augmentation
by: Li, Xiangju, et al.
Published: (2025)
by: Li, Xiangju, et al.
Published: (2025)
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
by: He, Linda, et al.
Published: (2025)
by: He, Linda, et al.
Published: (2025)
Rethinking Table Instruction Tuning
by: Deng, Naihao, et al.
Published: (2025)
by: Deng, Naihao, et al.
Published: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
Automatic Instruction Evolving for Large Language Models
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
by: Shu, Fan, et al.
Published: (2026)
by: Shu, Fan, et al.
Published: (2026)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
by: Ma, Huanhuan, et al.
Published: (2025)
by: Ma, Huanhuan, et al.
Published: (2025)
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
by: Latif, Ehsan, et al.
Published: (2023)
by: Latif, Ehsan, et al.
Published: (2023)
Integrating Expert Knowledge into Logical Programs via LLMs
by: Górski, Franciszek, et al.
Published: (2025)
by: Górski, Franciszek, et al.
Published: (2025)
Exploring Format Consistency for Instruction Tuning
by: Liang, Shihao, et al.
Published: (2023)
by: Liang, Shihao, et al.
Published: (2023)
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
by: Churina, Svetlana, et al.
Published: (2024)
by: Churina, Svetlana, et al.
Published: (2024)
Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER
by: Baroian, Andrei
Published: (2025)
by: Baroian, Andrei
Published: (2025)
The Instruction Gap: LLMs get lost in Following Instruction
by: Tripathi, Vishesh, et al.
Published: (2025)
by: Tripathi, Vishesh, et al.
Published: (2025)
Similar Items
-
Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency
by: Doostmohammadi, Ehsan, et al.
Published: (2025) -
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
by: Wang, Yidong, et al.
Published: (2023) -
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
by: Chen, Zhou, et al.
Published: (2025) -
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning
by: Zhou, Hang, et al.
Published: (2024) -
Fine-Tuning LLMs for Reliable Medical Question-Answering Services
by: Anaissi, Ali, et al.
Published: (2024)