Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Badshah, Sher, Sajjad, Hassan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
Large Language Models Report Subjective Experience Under Self-Referential Processing
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
von: Badshah, Sher, et al.
Veröffentlicht: (2026)
von: Badshah, Sher, et al.
Veröffentlicht: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
The Curious Case of In-Training Compression of State Space Models
von: Chahine, Makram, et al.
Veröffentlicht: (2025)
von: Chahine, Makram, et al.
Veröffentlicht: (2025)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
von: Plevris, Vagelis, et al.
Veröffentlicht: (2023)
von: Plevris, Vagelis, et al.
Veröffentlicht: (2023)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
Vibe-Creation: The Epistemology of Human-AI Emergent Cognition
von: Levin, Ilya
Veröffentlicht: (2026)
von: Levin, Ilya
Veröffentlicht: (2026)
PennyCoder: Efficient Domain-Specific LLMs for PennyLane-Based Quantum Code Generation
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
Reasoning Promotes Robustness in Theory of Mind Tasks
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
von: Rosales, Rafael, et al.
Veröffentlicht: (2025)
von: Rosales, Rafael, et al.
Veröffentlicht: (2025)
An Automatic Text Classification Method Based on Hierarchical Taxonomies, Neural Networks and Document Embedding: The NETHIC Tool
von: Lomasto, Luigi, et al.
Veröffentlicht: (2026)
von: Lomasto, Luigi, et al.
Veröffentlicht: (2026)
Mubeen AI: A Specialized Arabic Language Model for Heritage Preservation and User Intent Understanding
von: Aljafari, Mohammed, et al.
Veröffentlicht: (2025)
von: Aljafari, Mohammed, et al.
Veröffentlicht: (2025)
Stealth edits to large language models
von: Sutton, Oliver J., et al.
Veröffentlicht: (2024)
von: Sutton, Oliver J., et al.
Veröffentlicht: (2024)
Agentic AI Systems Applied to tasks in Financial Services: Modeling and model risk management crews
von: Okpala, Izunna, et al.
Veröffentlicht: (2025)
von: Okpala, Izunna, et al.
Veröffentlicht: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
CAPE: Corrective Actions from Precondition Errors using Large Language Models
von: Raman, Shreyas Sundara, et al.
Veröffentlicht: (2022)
von: Raman, Shreyas Sundara, et al.
Veröffentlicht: (2022)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
PLUGH: A Benchmark for Spatial Understanding and Reasoning in Large Language Models
von: Tikhonov, Alexey
Veröffentlicht: (2024)
von: Tikhonov, Alexey
Veröffentlicht: (2024)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
von: Xi, Wang, et al.
Veröffentlicht: (2025)
von: Xi, Wang, et al.
Veröffentlicht: (2025)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
GIM: Evaluating models via tasks that integrate multiple cognitive domains
von: Patel, Rohit, et al.
Veröffentlicht: (2026)
von: Patel, Rohit, et al.
Veröffentlicht: (2026)
Representing LLMs in Prompt Semantic Task Space
von: Kashani, Idan, et al.
Veröffentlicht: (2025)
von: Kashani, Idan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025) -
Large Language Models Report Subjective Experience Under Self-Referential Processing
von: Berg, Cameron, et al.
Veröffentlicht: (2025) -
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
von: Badshah, Sher, et al.
Veröffentlicht: (2026) -
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026) -
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)