How to Correctly Report LLM-as-a-Judge Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Chungpa, Zeng, Thomas, Jeong, Jongwon, Sohn, Jy-yong, Lee, Kangwook |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
Analysis of Using Sigmoid Loss for Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2024)
by: Lee, Chungpa, et al.
Published: (2024)
A Theoretical Framework for Preventing Class Collapse in Supervised Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2025)
by: Lee, Chungpa, et al.
Published: (2025)
On the Similarities of Embeddings in Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2025)
by: Lee, Chungpa, et al.
Published: (2025)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
by: Kim, Jungtaek, et al.
Published: (2026)
by: Kim, Jungtaek, et al.
Published: (2026)
Memorization Capacity for Additive Fine-Tuning with Small ReLU Networks
by: Sohn, Jy-yong, et al.
Published: (2024)
by: Sohn, Jy-yong, et al.
Published: (2024)
ERD: A Framework for Improving LLM Reasoning for Cognitive Distortion Classification
by: Lim, Sehee, et al.
Published: (2024)
by: Lim, Sehee, et al.
Published: (2024)
How to Choose a Threshold for an Evaluation Metric for Large Language Models
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
The Expressive Power of Low-Rank Adaptation
by: Zeng, Yuchen, et al.
Published: (2023)
by: Zeng, Yuchen, et al.
Published: (2023)
Detecting LLM-Generated Text with Performance Guarantees
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
Context-Alignment: Activating and Enhancing LLM Capabilities in Time Series
by: Hu, Yuxiao, et al.
Published: (2025)
by: Hu, Yuxiao, et al.
Published: (2025)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Bayesian Evaluation of Large Language Model Behavior
by: Longjohn, Rachel, et al.
Published: (2025)
by: Longjohn, Rachel, et al.
Published: (2025)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
by: Jeon, Jaebyeong, et al.
Published: (2025)
by: Jeon, Jaebyeong, et al.
Published: (2025)
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Efficient Inference for Noisy LLM-as-a-Judge Evaluation
by: Chen, Yiqun T, et al.
Published: (2026)
by: Chen, Yiqun T, et al.
Published: (2026)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
ENTP: Encoder-only Next Token Prediction
by: Ewer, Ethan, et al.
Published: (2024)
by: Ewer, Ethan, et al.
Published: (2024)
SigBERT: Combining Narrative Medical Reports and Rough Path Signature Theory for Survival Risk Estimation in Oncology
by: Minchella, Paul, et al.
Published: (2025)
by: Minchella, Paul, et al.
Published: (2025)
Buffer-based Gradient Projection for Continual Federated Learning
by: Dai, Shenghong, et al.
Published: (2024)
by: Dai, Shenghong, et al.
Published: (2024)
Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
by: Landesberg, Eddie, et al.
Published: (2025)
by: Landesberg, Eddie, et al.
Published: (2025)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
by: Kwon, Jihoon, et al.
Published: (2026)
by: Kwon, Jihoon, et al.
Published: (2026)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Classification errors distort findings in automated speech processing: examples and solutions from child-development research
by: Gautheron, Lucas, et al.
Published: (2025)
by: Gautheron, Lucas, et al.
Published: (2025)
Subjective Perspectives within Learned Representations Predict High-Impact Innovation
by: Cao, Likun, et al.
Published: (2025)
by: Cao, Likun, et al.
Published: (2025)
A meta-analysis on the performance of machine-learning based language models for sentiment analysis
by: Rohde, Elena, et al.
Published: (2025)
by: Rohde, Elena, et al.
Published: (2025)
Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
Ensemble Kalman filter for uncertainty in human language comprehension
by: Bhandari, Diksha, et al.
Published: (2025)
by: Bhandari, Diksha, et al.
Published: (2025)
Extracting Emotion Phrases from Tweets using BART
by: Rezapour, Mahdi
Published: (2024)
by: Rezapour, Mahdi
Published: (2024)
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
by: Li, Baishi, et al.
Published: (2026)
by: Li, Baishi, et al.
Published: (2026)
Dynamic Topic Language Model on Heterogeneous Children's Mental Health Clinical Notes
by: Ye, Hanwen, et al.
Published: (2023)
by: Ye, Hanwen, et al.
Published: (2023)
Specific language impairment (SLI) detection pipeline from transcriptions of spontaneous narratives
by: Arena, Santiago, et al.
Published: (2024)
by: Arena, Santiago, et al.
Published: (2024)
Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments
by: Imai, Kosuke, et al.
Published: (2024)
by: Imai, Kosuke, et al.
Published: (2024)
Explainable Automatic Grading with Neural Additive Models
by: Condor, Aubrey, et al.
Published: (2024)
by: Condor, Aubrey, et al.
Published: (2024)
Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective
by: Kim, Joonkyu, et al.
Published: (2025)
by: Kim, Joonkyu, et al.
Published: (2025)
Deep literature reviews: an application of fine-tuned language models to migration research
by: Iacus, Stefano M., et al.
Published: (2025)
by: Iacus, Stefano M., et al.
Published: (2025)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
by: Tierney, Graham, et al.
Published: (2025)
by: Tierney, Graham, et al.
Published: (2025)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
Similar Items
-
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026) -
Analysis of Using Sigmoid Loss for Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2024) -
A Theoretical Framework for Preventing Class Collapse in Supervised Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2025) -
On the Similarities of Embeddings in Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2025) -
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
by: Kim, Jungtaek, et al.
Published: (2026)