Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Meshkov, Aleksandr |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025)
von: Han, Steve, et al.
Veröffentlicht: (2025)
Evaluating ChatGPT as a Recommender System: A Rigorous Approach
von: Di Palma, Dario, et al.
Veröffentlicht: (2023)
von: Di Palma, Dario, et al.
Veröffentlicht: (2023)
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
von: Shu, Yiheng, et al.
Veröffentlicht: (2026)
von: Shu, Yiheng, et al.
Veröffentlicht: (2026)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
von: Zhou, Lexin, et al.
Veröffentlicht: (2025)
von: Zhou, Lexin, et al.
Veröffentlicht: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator
von: Sordo, Alessio, et al.
Veröffentlicht: (2026)
von: Sordo, Alessio, et al.
Veröffentlicht: (2026)
Adaptive Monitoring and Real-World Evaluation of Agentic AI Systems
von: Shukla, Manish
Veröffentlicht: (2025)
von: Shukla, Manish
Veröffentlicht: (2025)
Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
Call for Rigor in Reporting Quality of Instruction Tuning Data
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection
von: Gashkov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Gashkov, Aleksandr, et al.
Veröffentlicht: (2025)
TATRA: Training-Free Instance-Adaptive Prompting Through Rephrasing and Aggregation
von: Dziuba, Bartosz, et al.
Veröffentlicht: (2026)
von: Dziuba, Bartosz, et al.
Veröffentlicht: (2026)
Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
von: Zollo, Thomas P., et al.
Veröffentlicht: (2023)
von: Zollo, Thomas P., et al.
Veröffentlicht: (2023)
AI-Assisted Systematization for Evaluating GenAI Systems
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2026)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2026)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
Evaluation Framework for AI Systems in "the Wild"
von: Jabbour, Sarah, et al.
Veröffentlicht: (2025)
von: Jabbour, Sarah, et al.
Veröffentlicht: (2025)
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control
von: Liu, Huanshuo, et al.
Veröffentlicht: (2024)
von: Liu, Huanshuo, et al.
Veröffentlicht: (2024)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
von: Yao, Jing, et al.
Veröffentlicht: (2024)
von: Yao, Jing, et al.
Veröffentlicht: (2024)
FedDTRE: Federated Dialogue Generation Models Powered by Trustworthiness Evaluation
von: Lu, Shule, et al.
Veröffentlicht: (2025)
von: Lu, Shule, et al.
Veröffentlicht: (2025)
Knowledge Fusion via Bidirectional Information Aggregation
von: Zhai, Songlin, et al.
Veröffentlicht: (2025)
von: Zhai, Songlin, et al.
Veröffentlicht: (2025)
ArabIcros: AI-Powered Arabic Crossword Puzzle Generation for Educational Applications
von: Zeinalipour, Kamyar, et al.
Veröffentlicht: (2023)
von: Zeinalipour, Kamyar, et al.
Veröffentlicht: (2023)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
von: Garg, Madhav Krishan, et al.
Veröffentlicht: (2025)
von: Garg, Madhav Krishan, et al.
Veröffentlicht: (2025)
CURATRON: Complete and Robust Preference Data for Rigorous Alignment of Large Language Models
von: Nguyen, Son The, et al.
Veröffentlicht: (2024)
von: Nguyen, Son The, et al.
Veröffentlicht: (2024)
Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF
von: Fang, Yuan, et al.
Veröffentlicht: (2026)
von: Fang, Yuan, et al.
Veröffentlicht: (2026)
Towards a Science of Collective AI: LLM-based Multi-Agent Systems Need a Transition from Blind Trial-and-Error to Rigorous Science
von: Fan, Jingru, et al.
Veröffentlicht: (2026)
von: Fan, Jingru, et al.
Veröffentlicht: (2026)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform
von: Moses, Movina, et al.
Veröffentlicht: (2025)
von: Moses, Movina, et al.
Veröffentlicht: (2025)
Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summaries
von: Perez, Natalie, et al.
Veröffentlicht: (2026)
von: Perez, Natalie, et al.
Veröffentlicht: (2026)
Leveraging the Power of Large Language Models in Entity Linking via Adaptive Routing and Targeted Reasoning
von: Li, Yajie, et al.
Veröffentlicht: (2025)
von: Li, Yajie, et al.
Veröffentlicht: (2025)
Meanings and Feelings of Large Language Models: Observability of Latent States in Generative AI
von: Liu, Tian Yu, et al.
Veröffentlicht: (2024)
von: Liu, Tian Yu, et al.
Veröffentlicht: (2024)
From Phonemes to Meaning: Evaluating Large Language Models on Tamil
von: Varsha, Jeyarajalingam, et al.
Veröffentlicht: (2025)
von: Varsha, Jeyarajalingam, et al.
Veröffentlicht: (2025)
A Risk Ontology for Evaluating AI-Powered Psychotherapy Virtual Agents
von: Steenstra, Ian, et al.
Veröffentlicht: (2025)
von: Steenstra, Ian, et al.
Veröffentlicht: (2025)
Automatic Prompt Generation via Adaptive Selection of Prompting Techniques
von: Ikenoue, Yohei, et al.
Veröffentlicht: (2025)
von: Ikenoue, Yohei, et al.
Veröffentlicht: (2025)
AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
von: Huang, Shijue, et al.
Veröffentlicht: (2025)
von: Huang, Shijue, et al.
Veröffentlicht: (2025)
Adaptive Uncertainty Quantification for Generative AI
von: Kim, Jungeum, et al.
Veröffentlicht: (2024)
von: Kim, Jungeum, et al.
Veröffentlicht: (2024)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
What Generative Artificial Intelligence Means for Terminological Definitions
von: Martín, Antonio San
Veröffentlicht: (2024)
von: Martín, Antonio San
Veröffentlicht: (2024)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
von: Kostić, Bogdan, et al.
Veröffentlicht: (2026)
von: Kostić, Bogdan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
von: Liu, Yuhan, et al.
Veröffentlicht: (2025) -
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025) -
Evaluating ChatGPT as a Recommender System: A Rigorous Approach
von: Di Palma, Dario, et al.
Veröffentlicht: (2023) -
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
von: Shu, Yiheng, et al.
Veröffentlicht: (2026) -
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
von: Zhou, Lexin, et al.
Veröffentlicht: (2025)