Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Cheng, Yin, Haiyan, Tsang, Ivor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Effectiveness and Interpretability of Texts in LLM-based Time Series Models
by: Sun, Zhengke, et al.
Published: (2025)
by: Sun, Zhengke, et al.
Published: (2025)
HC$^2$L: Hybrid and Cooperative Contrastive Learning for Cross-lingual Spoken Language Understanding
by: Xing, Bowen, et al.
Published: (2024)
by: Xing, Bowen, et al.
Published: (2024)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
by: Dai, Hui, et al.
Published: (2024)
by: Dai, Hui, et al.
Published: (2024)
Time-Reversal Provides Unsupervised Feedback to LLMs
by: Varun, Yerram, et al.
Published: (2024)
by: Varun, Yerram, et al.
Published: (2024)
Collaborative Knowledge Infusion for Low-resource Stance Detection
by: Yan, Ming, et al.
Published: (2024)
by: Yan, Ming, et al.
Published: (2024)
Evaluating Role-Consistency in LLMs for Counselor Training
by: Rudolph, Eric, et al.
Published: (2026)
by: Rudolph, Eric, et al.
Published: (2026)
Unsupervised Layer-Wise Dynamic Test Time Adaptation for LLMs
by: Xu, Longhuan, et al.
Published: (2026)
by: Xu, Longhuan, et al.
Published: (2026)
Evaluating the Consistency of LLM Evaluators
by: Lee, Noah, et al.
Published: (2024)
by: Lee, Noah, et al.
Published: (2024)
Evaluating Large Language Models as Expert Annotators
by: Tseng, Yu-Min, et al.
Published: (2025)
by: Tseng, Yu-Min, et al.
Published: (2025)
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
by: Kowalyshyn, Katharine, et al.
Published: (2025)
by: Kowalyshyn, Katharine, et al.
Published: (2025)
Retrieve, Annotate, Evaluate, Repeat: Leveraging Multimodal LLMs for Large-Scale Product Retrieval Evaluation
by: Hosseini, Kasra, et al.
Published: (2024)
by: Hosseini, Kasra, et al.
Published: (2024)
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
Assessing Evaluation Metrics for Neural Test Oracle Generation
by: Shin, Jiho, et al.
Published: (2023)
by: Shin, Jiho, et al.
Published: (2023)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
Oracle-Checker Scheme for Evaluating a Generative Large Language Model
by: Zeng, Yueling Jenny, et al.
Published: (2024)
by: Zeng, Yueling Jenny, et al.
Published: (2024)
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
by: Zhang, Yidan, et al.
Published: (2024)
by: Zhang, Yidan, et al.
Published: (2024)
AXCEL: Automated eXplainable Consistency Evaluation using LLMs
by: Sreekar, P Aditya, et al.
Published: (2024)
by: Sreekar, P Aditya, et al.
Published: (2024)
Dynamic Evaluation for Oversensitivity in LLMs
by: Pu, Sophia Xiao, et al.
Published: (2025)
by: Pu, Sophia Xiao, et al.
Published: (2025)
Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
by: Alakeel, Yara, et al.
Published: (2026)
by: Alakeel, Yara, et al.
Published: (2026)
Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
by: Hejabi, Parsa, et al.
Published: (2025)
by: Hejabi, Parsa, et al.
Published: (2025)
Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images
by: Rykov, Elisei, et al.
Published: (2025)
by: Rykov, Elisei, et al.
Published: (2025)
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
by: Rashkin, Hannah, et al.
Published: (2025)
by: Rashkin, Hannah, et al.
Published: (2025)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
by: Wang, Xingyao, et al.
Published: (2023)
by: Wang, Xingyao, et al.
Published: (2023)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Through the Prism of Culture: Evaluating LLMs' Understanding of Indian Subcultures and Traditions
by: Chhikara, Garima, et al.
Published: (2025)
by: Chhikara, Garima, et al.
Published: (2025)
Unsupervised Contrast-Consistent Ranking with Language Models
by: Stoehr, Niklas, et al.
Published: (2023)
by: Stoehr, Niklas, et al.
Published: (2023)
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
by: Hu, Gang, et al.
Published: (2026)
by: Hu, Gang, et al.
Published: (2026)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024)
by: Lee, Jaehyeok, et al.
Published: (2024)
The Biased Oracle: Assessing LLMs' Understandability and Empathy in Medical Diagnoses
by: Yao, Jianzhou, et al.
Published: (2025)
by: Yao, Jianzhou, et al.
Published: (2025)
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
by: Xia, Congying, et al.
Published: (2024)
by: Xia, Congying, et al.
Published: (2024)
Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
by: Tan, Hexiang, et al.
Published: (2025)
by: Tan, Hexiang, et al.
Published: (2025)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Code Aesthetics with Agentic Reward Feedback
by: Xiao, Bang, et al.
Published: (2025)
by: Xiao, Bang, et al.
Published: (2025)
Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications
by: Megahed, Fadel M., et al.
Published: (2025)
by: Megahed, Fadel M., et al.
Published: (2025)
LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks
by: Wan, Jiayong, et al.
Published: (2026)
by: Wan, Jiayong, et al.
Published: (2026)
Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation
by: Wu, Sophie, et al.
Published: (2026)
by: Wu, Sophie, et al.
Published: (2026)
The Agentic Leash: Extracting Causal Feedback Fuzzy Cognitive Maps with LLMs
by: Panda, Akash Kumar, et al.
Published: (2025)
by: Panda, Akash Kumar, et al.
Published: (2025)
BaseCal: Unsupervised Confidence Calibration via Base Model Signals
by: Tan, Hexiang, et al.
Published: (2026)
by: Tan, Hexiang, et al.
Published: (2026)
Similar Items
-
Exploring the Effectiveness and Interpretability of Texts in LLM-based Time Series Models
by: Sun, Zhengke, et al.
Published: (2025) -
HC$^2$L: Hybrid and Cooperative Contrastive Learning for Cross-lingual Spoken Language Understanding
by: Xing, Bowen, et al.
Published: (2024) -
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025) -
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
by: Dai, Hui, et al.
Published: (2024) -
Time-Reversal Provides Unsupervised Feedback to LLMs
by: Varun, Yerram, et al.
Published: (2024)