From Guidelines to Guarantees: A Graph-Based Evaluation Harness for Domain-Specific Evaluation of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lundin, Jessica M., Nakakana, Usman Nasir, Chabot-Couture, Guillaume |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
by: Alhanai, Tuka, et al.
Published: (2024)
by: Alhanai, Tuka, et al.
Published: (2024)
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
by: Mustafa, Akram, et al.
Published: (2025)
by: Mustafa, Akram, et al.
Published: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
by: Xiao, Yilin, et al.
Published: (2025)
by: Xiao, Yilin, et al.
Published: (2025)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
by: Seo, Jean, et al.
Published: (2026)
by: Seo, Jean, et al.
Published: (2026)
Qworld: Question-Specific Evaluation Criteria for LLMs
by: Gao, Shanghua, et al.
Published: (2026)
by: Gao, Shanghua, et al.
Published: (2026)
Evaluation Ethics of LLMs in Legal Domain
by: Zhang, Ruizhe, et al.
Published: (2024)
by: Zhang, Ruizhe, et al.
Published: (2024)
Evaluating ChatGPT on Nuclear Domain-Specific Data
by: Anwar, Muhammad, et al.
Published: (2024)
by: Anwar, Muhammad, et al.
Published: (2024)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
by: Nadeem, Afrozah, et al.
Published: (2026)
by: Nadeem, Afrozah, et al.
Published: (2026)
Cardiverse: Harnessing LLMs for Novel Card Game Prototyping
by: Li, Danrui, et al.
Published: (2025)
by: Li, Danrui, et al.
Published: (2025)
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation
by: Zheng, Huawei, et al.
Published: (2026)
by: Zheng, Huawei, et al.
Published: (2026)
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
by: Gao, Joshua, et al.
Published: (2025)
by: Gao, Joshua, et al.
Published: (2025)
No Text Needed: Forecasting MT Quality and Inequity from Fertility and Metadata
by: Lundin, Jessica M., et al.
Published: (2025)
by: Lundin, Jessica M., et al.
Published: (2025)
From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs
by: Chen, Yingjian, et al.
Published: (2026)
by: Chen, Yingjian, et al.
Published: (2026)
LLMs to Support a Domain Specific Knowledge Assistant
by: Lovin, Maria-Flavia
Published: (2025)
by: Lovin, Maria-Flavia
Published: (2025)
A Decade-Scale Benchmark Evaluating LLMs' Clinical Practice Guidelines Detection and Adherence in Multi-turn Conversations
by: Tan, Andong, et al.
Published: (2026)
by: Tan, Andong, et al.
Published: (2026)
Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval
by: Bhavsar, Vivek, et al.
Published: (2025)
by: Bhavsar, Vivek, et al.
Published: (2025)
All-in-One Tuning and Structural Pruning for Domain-Specific LLMs
by: Lu, Lei, et al.
Published: (2024)
by: Lu, Lei, et al.
Published: (2024)
PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs
by: Liu, An, et al.
Published: (2024)
by: Liu, An, et al.
Published: (2024)
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
by: Pei, Aihua, et al.
Published: (2024)
by: Pei, Aihua, et al.
Published: (2024)
Adapting LLMs for the Medical Domain in Portuguese: A Study on Fine-Tuning and Model Evaluation
by: Paiola, Pedro Henrique, et al.
Published: (2024)
by: Paiola, Pedro Henrique, et al.
Published: (2024)
Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs
by: Bouchard, Dylan
Published: (2024)
by: Bouchard, Dylan
Published: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
A Comprehensive Evaluation of Cognitive Biases in LLMs
by: Malberg, Simon, et al.
Published: (2024)
by: Malberg, Simon, et al.
Published: (2024)
AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM Chatbots
by: Zhao, Xinjie, et al.
Published: (2025)
by: Zhao, Xinjie, et al.
Published: (2025)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
by: Rahman, A B M Ashikur, et al.
Published: (2024)
by: Rahman, A B M Ashikur, et al.
Published: (2024)
Automated Evaluation of Classroom Instructional Support with LLMs and BoWs: Connecting Global Predictions to Specific Feedback
by: Whitehill, Jacob, et al.
Published: (2023)
by: Whitehill, Jacob, et al.
Published: (2023)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
by: Zhou, Chengliang, et al.
Published: (2025)
by: Zhou, Chengliang, et al.
Published: (2025)
KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs
by: Pei, Aihua, et al.
Published: (2024)
by: Pei, Aihua, et al.
Published: (2024)
Data Analysis and Performance Evaluation of Simulation Deduction Based on LLMs
by: Zhang, Shansi, et al.
Published: (2025)
by: Zhang, Shansi, et al.
Published: (2025)
Agentic Adversarial QA for Improving Domain-Specific LLMs
by: Grari, Vincent, et al.
Published: (2026)
by: Grari, Vincent, et al.
Published: (2026)
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
by: Sun, Chongyan, et al.
Published: (2024)
by: Sun, Chongyan, et al.
Published: (2024)
Knowledge-Graph Based RAG System Evaluation Framework
by: Dong, Sicheng, et al.
Published: (2025)
by: Dong, Sicheng, et al.
Published: (2025)
GameTraversalBenchmark: Evaluating Planning Abilities Of Large Language Models Through Traversing 2D Game Maps
by: Nasir, Muhammad Umair, et al.
Published: (2024)
by: Nasir, Muhammad Umair, et al.
Published: (2024)
Comparative Evaluation of ChatGPT and DeepSeek Across Key NLP Tasks: Strengths, Weaknesses, and Domain-Specific Performance
by: Etaiwi, Wael, et al.
Published: (2025)
by: Etaiwi, Wael, et al.
Published: (2025)
Multi-Domain ABSA Conversation Dataset Generation via LLMs for Real-World Evaluation and Model Comparison
by: Pandit, Tejul, et al.
Published: (2025)
by: Pandit, Tejul, et al.
Published: (2025)
Evaluating Compliance with Visualization Guidelines in Diagrams for Scientific Publications Using Large Vision Language Models
by: Rückert, Johannes, et al.
Published: (2025)
by: Rückert, Johannes, et al.
Published: (2025)
Similar Items
-
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
by: Alhanai, Tuka, et al.
Published: (2024) -
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
by: Mustafa, Akram, et al.
Published: (2025) -
Fairness Evaluation and Inference Level Mitigation in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025) -
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
by: Xiao, Yilin, et al.
Published: (2025) -
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
by: Seo, Jean, et al.
Published: (2026)