TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Badshah, Sher, Emami, Ali, Sajjad, Hassan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
von: Badshah, Sher, et al.
Veröffentlicht: (2026)
von: Badshah, Sher, et al.
Veröffentlicht: (2026)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
von: Budagam, Devichand, et al.
Veröffentlicht: (2024)
von: Budagam, Devichand, et al.
Veröffentlicht: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
Accurate and Energy Efficient: Local Retrieval-Augmented Generation Models Outperform Commercial Large Language Models in Medical Tasks
von: Vrettos, Konstantinos, et al.
Veröffentlicht: (2025)
von: Vrettos, Konstantinos, et al.
Veröffentlicht: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
von: Kytöniemi, Joona, et al.
Veröffentlicht: (2025)
von: Kytöniemi, Joona, et al.
Veröffentlicht: (2025)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
von: Sandan, Isik Baran, et al.
Veröffentlicht: (2025)
von: Sandan, Isik Baran, et al.
Veröffentlicht: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Robustness of Large Language Models to Perturbations in Text
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
A comprehensive taxonomy of hallucinations in Large Language Models
von: Cossio, Manuel
Veröffentlicht: (2025)
von: Cossio, Manuel
Veröffentlicht: (2025)
Evaluating Class Membership Relations in Knowledge Graphs using Large Language Models
von: Allen, Bradley P., et al.
Veröffentlicht: (2024)
von: Allen, Bradley P., et al.
Veröffentlicht: (2024)
PatentGPT: A Large Language Model for Intellectual Property
von: Bai, Zilong, et al.
Veröffentlicht: (2024)
von: Bai, Zilong, et al.
Veröffentlicht: (2024)
Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models
von: Niu, Shuai, et al.
Veröffentlicht: (2024)
von: Niu, Shuai, et al.
Veröffentlicht: (2024)
Inference to the Best Explanation in Large Language Models
von: Dalal, Dhairya, et al.
Veröffentlicht: (2024)
von: Dalal, Dhairya, et al.
Veröffentlicht: (2024)
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
von: Gu, Yu, et al.
Veröffentlicht: (2024)
von: Gu, Yu, et al.
Veröffentlicht: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
Streamlining Redundant Layers to Compress Large Language Models
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
von: Chang, Edward Y.
Veröffentlicht: (2024)
von: Chang, Edward Y.
Veröffentlicht: (2024)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
von: Hawkins, John
Veröffentlicht: (2025)
von: Hawkins, John
Veröffentlicht: (2025)
Bielik 11B v3: Multilingual Large Language Model for European Languages
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models
von: Lin, Tianqianjin, et al.
Veröffentlicht: (2025)
von: Lin, Tianqianjin, et al.
Veröffentlicht: (2025)
Prompt-Time Symbolic Knowledge Capture with Large Language Models
von: Çöplü, Tolga, et al.
Veröffentlicht: (2024)
von: Çöplü, Tolga, et al.
Veröffentlicht: (2024)
Artificial Phantasia: Emergent Mental Imagery in Large Language Models
von: McCarty, Morgan, et al.
Veröffentlicht: (2025)
von: McCarty, Morgan, et al.
Veröffentlicht: (2025)
Argumentative Large Language Models for Explainable and Contestable Claim Verification
von: Freedman, Gabriel, et al.
Veröffentlicht: (2024)
von: Freedman, Gabriel, et al.
Veröffentlicht: (2024)
Reasoning over Uncertain Text by Generative Large Language Models
von: Nafar, Aliakbar, et al.
Veröffentlicht: (2024)
von: Nafar, Aliakbar, et al.
Veröffentlicht: (2024)
Multi-Turn Interactions for Text-to-SQL with Large Language Models
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
von: Paech, Samuel J.
Veröffentlicht: (2023)
von: Paech, Samuel J.
Veröffentlicht: (2023)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
von: Benkirane, Kenza, et al.
Veröffentlicht: (2024)
von: Benkirane, Kenza, et al.
Veröffentlicht: (2024)
Search-R3: Unifying Reasoning and Embedding in Large Language Models
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
Can Large Language Models perform Relation-based Argument Mining?
von: Gorur, Deniz, et al.
Veröffentlicht: (2024)
von: Gorur, Deniz, et al.
Veröffentlicht: (2024)
EMNLP: Educator-role Moral and Normative Large Language Models Profiling
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
HAMMER: Hamiltonian Curiosity Augmented Large Language Model Reinforcement
von: Yang, Ming, et al.
Veröffentlicht: (2025)
von: Yang, Ming, et al.
Veröffentlicht: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Attentive Reasoning Queries: A Systematic Method for Optimizing Instruction-Following in Large Language Models
von: Karov, Bar, et al.
Veröffentlicht: (2025)
von: Karov, Bar, et al.
Veröffentlicht: (2025)
Internal Reasoning vs. External Control: A Thermodynamic Analysis of Sycophancy in Large Language Models
von: Chang, Edward Y.
Veröffentlicht: (2025)
von: Chang, Edward Y.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024) -
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
von: Badshah, Sher, et al.
Veröffentlicht: (2026) -
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025) -
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
von: Budagam, Devichand, et al.
Veröffentlicht: (2024)