The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Abeysinghe, Bhashithe, Circi, Ruhan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
par: De Michele, William, et autres
Publié: (2025)
par: De Michele, William, et autres
Publié: (2025)
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
par: Refai, Dania, et autres
Publié: (2025)
par: Refai, Dania, et autres
Publié: (2025)
Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications
par: Kogan, David, et autres
Publié: (2025)
par: Kogan, David, et autres
Publié: (2025)
LLM Prompt Evaluation for Educational Applications
par: Holmes, Langdon, et autres
Publié: (2026)
par: Holmes, Langdon, et autres
Publié: (2026)
Rethinking Human Preference Evaluation of LLM Rationales
par: Li, Ziang, et autres
Publié: (2025)
par: Li, Ziang, et autres
Publié: (2025)
An LLM-Based Approach for Insight Generation in Data Analysis
par: Pérez, Alberto Sánchez, et autres
Publié: (2025)
par: Pérez, Alberto Sánchez, et autres
Publié: (2025)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
par: Cho, Yousang, et autres
Publié: (2025)
par: Cho, Yousang, et autres
Publié: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
par: Yehudai, Asaf, et autres
Publié: (2026)
par: Yehudai, Asaf, et autres
Publié: (2026)
DIALEVAL: Automated Type-Theoretic Evaluation of LLM Instruction Following
par: Basta, Nardine, et autres
Publié: (2026)
par: Basta, Nardine, et autres
Publié: (2026)
Automated Concept Discovery for LLM-as-a-Judge Preference Analysis
par: Wedgwood, James, et autres
Publié: (2026)
par: Wedgwood, James, et autres
Publié: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
par: Fayyaz, Mohsen, et autres
Publié: (2024)
par: Fayyaz, Mohsen, et autres
Publié: (2024)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
par: Atasoy, I. F., et autres
Publié: (2026)
par: Atasoy, I. F., et autres
Publié: (2026)
Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
par: Wang, Xi, et autres
Publié: (2025)
par: Wang, Xi, et autres
Publié: (2025)
Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
par: McGinness, Lachlan, et autres
Publié: (2024)
par: McGinness, Lachlan, et autres
Publié: (2024)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
par: Li, Yanhong, et autres
Publié: (2025)
par: Li, Yanhong, et autres
Publié: (2025)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
par: Sitaram, Sunayana, et autres
Publié: (2025)
par: Sitaram, Sunayana, et autres
Publié: (2025)
LITERA: An LLM Based Approach to Latin-to-English Translation
par: Rosu, Paul
Publié: (2025)
par: Rosu, Paul
Publié: (2025)
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
par: Richardson, Andrew Keenan, et autres
Publié: (2025)
par: Richardson, Andrew Keenan, et autres
Publié: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
par: Wu, JiaRu, et autres
Publié: (2025)
par: Wu, JiaRu, et autres
Publié: (2025)
Trustworthy LLM-Mediated Communication: Evaluating Information Fidelity in LLM as a Communicator (LAAC) Framework in Multiple Application Domains
par: Rafi, Mohammed Musthafa, et autres
Publié: (2025)
par: Rafi, Mohammed Musthafa, et autres
Publié: (2025)
Dissecting Human and LLM Preferences
par: Li, Junlong, et autres
Publié: (2024)
par: Li, Junlong, et autres
Publié: (2024)
Datarus-R1: An Adaptive Multi-Step Reasoning LLM for Automated Data Analysis
par: Chaliah, Ayoub Ben, et autres
Publié: (2025)
par: Chaliah, Ayoub Ben, et autres
Publié: (2025)
Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach
par: Gundawar, Atharva, et autres
Publié: (2024)
par: Gundawar, Atharva, et autres
Publié: (2024)
Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity
par: Liu, Qiawen Ella, et autres
Publié: (2026)
par: Liu, Qiawen Ella, et autres
Publié: (2026)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
par: Lu, Dongxu, et autres
Publié: (2025)
par: Lu, Dongxu, et autres
Publié: (2025)
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
par: Sun, Peng, et autres
Publié: (2026)
par: Sun, Peng, et autres
Publié: (2026)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
par: Arabzadeh, Negar, et autres
Publié: (2024)
par: Arabzadeh, Negar, et autres
Publié: (2024)
Evaluating Morphological Compositional Generalization in Large Language Models
par: Ismayilzada, Mete, et autres
Publié: (2024)
par: Ismayilzada, Mete, et autres
Publié: (2024)
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
par: Kanithi, Praveenkumar, et autres
Publié: (2024)
par: Kanithi, Praveenkumar, et autres
Publié: (2024)
GuideLLM: Exploring LLM-Guided Conversation with Applications in Autobiography Interviewing
par: Duan, Jinhao, et autres
Publié: (2025)
par: Duan, Jinhao, et autres
Publié: (2025)
LLM and Agent-Driven Data Analysis: A Systematic Approach for Enterprise Applications and System-level Deployment
par: Wang, Xi, et autres
Publié: (2025)
par: Wang, Xi, et autres
Publié: (2025)
Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study
par: Yamagishi, Yosuke, et autres
Publié: (2026)
par: Yamagishi, Yosuke, et autres
Publié: (2026)
From Internal Representations to Text Quality: A Geometric Approach to LLM Evaluation
par: Yusupov, Viacheslav, et autres
Publié: (2025)
par: Yusupov, Viacheslav, et autres
Publié: (2025)
Optimizing In-Context Demonstrations for LLM-based Automated Grading
par: Chu, Yucheng, et autres
Publié: (2026)
par: Chu, Yucheng, et autres
Publié: (2026)
Automated Survey Collection with LLM-based Conversational Agents
par: Kaiyrbekov, Kurmanbek, et autres
Publié: (2025)
par: Kaiyrbekov, Kurmanbek, et autres
Publié: (2025)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
par: Rahman, Mizanur, et autres
Publié: (2025)
par: Rahman, Mizanur, et autres
Publié: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
par: Kim, Eunsu, et autres
Publié: (2024)
par: Kim, Eunsu, et autres
Publié: (2024)
Investigating LLM Applications in E-Commerce
par: Palen-Michel, Chester, et autres
Publié: (2024)
par: Palen-Michel, Chester, et autres
Publié: (2024)
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator
par: Sordo, Alessio, et autres
Publié: (2026)
par: Sordo, Alessio, et autres
Publié: (2026)
Optimization Techniques for Sentiment Analysis Based on LLM (GPT-3)
par: Zhan, Tong, et autres
Publié: (2024)
par: Zhan, Tong, et autres
Publié: (2024)
Documents similaires
-
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
par: De Michele, William, et autres
Publié: (2025) -
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
par: Refai, Dania, et autres
Publié: (2025) -
Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications
par: Kogan, David, et autres
Publié: (2025) -
LLM Prompt Evaluation for Educational Applications
par: Holmes, Langdon, et autres
Publié: (2026) -
Rethinking Human Preference Evaluation of LLM Rationales
par: Li, Ziang, et autres
Publié: (2025)