Saved in:
| Main Authors: | Srun, Nalin, Rastin, Parisa, Cabanes, Guénaël, Assala, Lydia Boudjeloud |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.09624 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
by: Chu, Seongyeub, et al.
Published: (2026)
by: Chu, Seongyeub, et al.
Published: (2026)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
by: Rafid, Ahmed, et al.
Published: (2026)
by: Rafid, Ahmed, et al.
Published: (2026)
ReproHum #0087-01: Human Evaluation Reproduction Report for Generating Fact Checking Explanations
by: Loakman, Tyler, et al.
Published: (2024)
by: Loakman, Tyler, et al.
Published: (2024)
FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks
by: Premsri, Tanawan, et al.
Published: (2025)
by: Premsri, Tanawan, et al.
Published: (2025)
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models
by: Fuad, Kazi Ahmed Asif, et al.
Published: (2024)
by: Fuad, Kazi Ahmed Asif, et al.
Published: (2024)
MILE: A Mutation Testing Framework of In-Context Learning Systems
by: Wei, Zeming, et al.
Published: (2024)
by: Wei, Zeming, et al.
Published: (2024)
HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations
by: Sicilia, Anthony, et al.
Published: (2023)
by: Sicilia, Anthony, et al.
Published: (2023)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing
by: Premsri, Tanawan, et al.
Published: (2025)
by: Premsri, Tanawan, et al.
Published: (2025)
MindRef: Mimicking Human Memory for Hierarchical Reference Retrieval with Fine-Grained Location Awareness
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
by: Dejl, Adam, et al.
Published: (2026)
by: Dejl, Adam, et al.
Published: (2026)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
An optimal baseline selection methodology for data-driven damage detection and temperature compensation in acousto-ultrasonics
by: Torres-Arredondo, M-A, et al.
Published: (2025)
by: Torres-Arredondo, M-A, et al.
Published: (2025)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
by: Weck, Benno, et al.
Published: (2026)
by: Weck, Benno, et al.
Published: (2026)
Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
by: Broy, Frederik, et al.
Published: (2025)
by: Broy, Frederik, et al.
Published: (2025)
BotEval: Facilitating Interactive Human Evaluation
by: Cho, Hyundong, et al.
Published: (2024)
by: Cho, Hyundong, et al.
Published: (2024)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
by: Zhang, Qiyuan, et al.
Published: (2024)
by: Zhang, Qiyuan, et al.
Published: (2024)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
by: Hua, Yilun, et al.
Published: (2026)
by: Hua, Yilun, et al.
Published: (2026)
RepEval: Effective Text Evaluation with LLM Representation
by: Sheng, Shuqian, et al.
Published: (2024)
by: Sheng, Shuqian, et al.
Published: (2024)
BatchEval: Towards Human-like Text Evaluation
by: Yuan, Peiwen, et al.
Published: (2023)
by: Yuan, Peiwen, et al.
Published: (2023)
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
by: Liu, Yuxuan, et al.
Published: (2024)
by: Liu, Yuxuan, et al.
Published: (2024)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
RefChecker: Reference-based Fine-grained Hallucination Checker and Benchmark for Large Language Models
by: Hu, Xiangkun, et al.
Published: (2024)
by: Hu, Xiangkun, et al.
Published: (2024)
Reference-Free Evaluation of Taxonomies
by: Wullschleger, Pascal, et al.
Published: (2025)
by: Wullschleger, Pascal, et al.
Published: (2025)
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
by: Jiang, Lai, et al.
Published: (2025)
by: Jiang, Lai, et al.
Published: (2025)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
by: Alyahya, Hisham A., et al.
Published: (2025)
by: Alyahya, Hisham A., et al.
Published: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation
by: Shomee, Homaira Huda, et al.
Published: (2026)
by: Shomee, Homaira Huda, et al.
Published: (2026)
HumT DumT: Measuring and controlling human-like language in LLMs
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
Eval4Sim: An Evaluation Framework for Persona Simulation
by: Bao, Eliseo, et al.
Published: (2026)
by: Bao, Eliseo, et al.
Published: (2026)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
by: Liang, Chumeng, et al.
Published: (2025)
by: Liang, Chumeng, et al.
Published: (2025)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
by: Shen, Chengyu, et al.
Published: (2026)
by: Shen, Chengyu, et al.
Published: (2026)
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
by: Gritta, Milan, et al.
Published: (2024)
by: Gritta, Milan, et al.
Published: (2024)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
by: Sansford, Hannah, et al.
Published: (2024)
by: Sansford, Hannah, et al.
Published: (2024)
Exploring Spatial Language Grounding Through Referring Expressions
by: Tumu, Akshar, et al.
Published: (2025)
by: Tumu, Akshar, et al.
Published: (2025)
NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
Similar Items
-
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
by: Chu, Seongyeub, et al.
Published: (2026) -
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
by: Rafid, Ahmed, et al.
Published: (2026) -
ReproHum #0087-01: Human Evaluation Reproduction Report for Generating Fact Checking Explanations
by: Loakman, Tyler, et al.
Published: (2024) -
FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks
by: Premsri, Tanawan, et al.
Published: (2025) -
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models
by: Fuad, Kazi Ahmed Asif, et al.
Published: (2024)