Improving LLM Leaderboards with Psychometrical Methodology
Fuente:
arXiv
Guardado en:
| Autor principal: | Federiakin, Denis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
por: von der Heyde, Leah, et al.
Publicado: (2024)
por: von der Heyde, Leah, et al.
Publicado: (2024)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
por: Zhou, Cai, et al.
Publicado: (2026)
por: Zhou, Cai, et al.
Publicado: (2026)
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
por: Davoudi, Seyed Pouyan Mousavi, et al.
Publicado: (2025)
por: Davoudi, Seyed Pouyan Mousavi, et al.
Publicado: (2025)
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages
por: Schöffel, Matthias, et al.
Publicado: (2026)
por: Schöffel, Matthias, et al.
Publicado: (2026)
RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning
por: Jin, Congyun, et al.
Publicado: (2024)
por: Jin, Congyun, et al.
Publicado: (2024)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
por: Oleson, Jon
Publicado: (2024)
por: Oleson, Jon
Publicado: (2024)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
por: Devunuri, Saipraneeth, et al.
Publicado: (2024)
por: Devunuri, Saipraneeth, et al.
Publicado: (2024)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
por: Hardy, Michael
Publicado: (2024)
por: Hardy, Michael
Publicado: (2024)
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
por: Hu, Zhengmian, et al.
Publicado: (2024)
por: Hu, Zhengmian, et al.
Publicado: (2024)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
por: Nie, Allen, et al.
Publicado: (2024)
por: Nie, Allen, et al.
Publicado: (2024)
Metacognitive Myopia in Large Language Models
por: Scholten, Florian, et al.
Publicado: (2024)
por: Scholten, Florian, et al.
Publicado: (2024)
League: Leaderboard Generation on Demand
por: Wu, Jian, et al.
Publicado: (2025)
por: Wu, Jian, et al.
Publicado: (2025)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
por: Park, Chanjun, et al.
Publicado: (2024)
por: Park, Chanjun, et al.
Publicado: (2024)
Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
por: Kokkodis, Marios, et al.
Publicado: (2025)
por: Kokkodis, Marios, et al.
Publicado: (2025)
ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution
por: de Carvalho, Gonçalo Hora, et al.
Publicado: (2025)
por: de Carvalho, Gonçalo Hora, et al.
Publicado: (2025)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
Language-Dependent Political Bias in AI: A Study of ChatGPT and Gemini
por: Yuksel, Dogus, et al.
Publicado: (2025)
por: Yuksel, Dogus, et al.
Publicado: (2025)
Domain-Shift-Aware Conformal Prediction for Large Language Models
por: Lin, Zhexiao, et al.
Publicado: (2025)
por: Lin, Zhexiao, et al.
Publicado: (2025)
Reliable and Efficient Amortized Model-based Evaluation
por: Truong, Sang, et al.
Publicado: (2025)
por: Truong, Sang, et al.
Publicado: (2025)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
por: Jantre, Sanket, et al.
Publicado: (2025)
por: Jantre, Sanket, et al.
Publicado: (2025)
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
por: Dobbins, Nic, et al.
Publicado: (2025)
por: Dobbins, Nic, et al.
Publicado: (2025)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
por: Kuzmanko, Jonathan
Publicado: (2025)
por: Kuzmanko, Jonathan
Publicado: (2025)
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
por: Yuan, Hongyi, et al.
Publicado: (2024)
por: Yuan, Hongyi, et al.
Publicado: (2024)
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
por: Tschisgale, Paul, et al.
Publicado: (2026)
por: Tschisgale, Paul, et al.
Publicado: (2026)
A Rational Analysis of the Speech-to-Song Illusion
por: Marjieh, Raja, et al.
Publicado: (2024)
por: Marjieh, Raja, et al.
Publicado: (2024)
Augmented Risk Prediction for the Onset of Alzheimer's Disease from Electronic Health Records with Large Language Models
por: Wang, Jiankun, et al.
Publicado: (2024)
por: Wang, Jiankun, et al.
Publicado: (2024)
Crowdsourced Adaptive Surveys
por: Velez, Yamil
Publicado: (2024)
por: Velez, Yamil
Publicado: (2024)
Limits of Large Language Models in Debating Humans
por: Flamino, James, et al.
Publicado: (2024)
por: Flamino, James, et al.
Publicado: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
por: Kim, Hyeonwoo, et al.
Publicado: (2024)
por: Kim, Hyeonwoo, et al.
Publicado: (2024)
Exploring the Latest LLMs for Leaderboard Extraction
por: Kabongo, Salomon, et al.
Publicado: (2024)
por: Kabongo, Salomon, et al.
Publicado: (2024)
ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms
por: Vail, Alexandria K., et al.
Publicado: (2026)
por: Vail, Alexandria K., et al.
Publicado: (2026)
The Leaderboard Illusion
por: Singh, Shivalika, et al.
Publicado: (2025)
por: Singh, Shivalika, et al.
Publicado: (2025)
Removing Spurious Correlation from Neural Network Interpretations
por: Fotouhi, Milad, et al.
Publicado: (2024)
por: Fotouhi, Milad, et al.
Publicado: (2024)
Language Models as Causal Effect Generators
por: Bynum, Lucius E. J., et al.
Publicado: (2024)
por: Bynum, Lucius E. J., et al.
Publicado: (2024)
Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
por: Hang, Haotian, et al.
Publicado: (2025)
por: Hang, Haotian, et al.
Publicado: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
por: Park, Chanjun, et al.
Publicado: (2024)
por: Park, Chanjun, et al.
Publicado: (2024)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
por: Myrzakhan, Aidar, et al.
Publicado: (2024)
por: Myrzakhan, Aidar, et al.
Publicado: (2024)
Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational Theory
por: Yuan, Yunhao, et al.
Publicado: (2025)
por: Yuan, Yunhao, et al.
Publicado: (2025)
Ejemplares similares
-
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
por: Wang, Hao, et al.
Publicado: (2026) -
United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
por: von der Heyde, Leah, et al.
Publicado: (2024) -
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
por: Zhou, Cai, et al.
Publicado: (2026) -
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
por: Davoudi, Seyed Pouyan Mousavi, et al.
Publicado: (2025) -
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages
por: Schöffel, Matthias, et al.
Publicado: (2026)