A Position Paper on the Automatic Generation of Machine Learning Leaderboards
Fuente:
arXiv
Saved in:
| Main Authors: | Timmer, Roelien C, Hou, Yufang, Wan, Stephen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaLead: A Comprehensive Human-Curated Leaderboard Dataset for Transparent Reporting of Machine Learning Experiments
by: Timmer, Roelien C., et al.
Published: (2026)
by: Timmer, Roelien C., et al.
Published: (2026)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
by: Timmer, Roelien C., et al.
Published: (2024)
by: Timmer, Roelien C., et al.
Published: (2024)
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
by: Şahinuç, Furkan, et al.
Published: (2024)
by: Şahinuç, Furkan, et al.
Published: (2024)
Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
by: Lyu, Y., et al.
Published: (2025)
by: Lyu, Y., et al.
Published: (2025)
Exploring the Potential Role of Generative AI in the TRAPD Procedure for Survey Translation
by: Metheney, Erica Ann, et al.
Published: (2024)
by: Metheney, Erica Ann, et al.
Published: (2024)
Improving Probabilistic Models in Text Classification via Active Learning
by: Bosley, Mitchell, et al.
Published: (2022)
by: Bosley, Mitchell, et al.
Published: (2022)
A Finite-Calibration Regime Map for LLM Judge Panels
by: Zhu, Bin, et al.
Published: (2026)
by: Zhu, Bin, et al.
Published: (2026)
A New Semisupervised Technique for Polarity Analysis using Masked Language Models
by: Watanabe, Kohei
Published: (2026)
by: Watanabe, Kohei
Published: (2026)
Distributed Asymmetric Allocation: A Topic Model for Large Imbalanced Corpora in Social Sciences
by: Watanabe, Kohei
Published: (2025)
by: Watanabe, Kohei
Published: (2025)
Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study
by: Okada, Kensuke, et al.
Published: (2026)
by: Okada, Kensuke, et al.
Published: (2026)
A chart review process aided by natural language processing and multi-wave adaptive sampling to expedite validation of code-based algorithms for large database studies
by: Wang, Shirley V, et al.
Published: (2025)
by: Wang, Shirley V, et al.
Published: (2025)
The Nature of NLP: Analyzing Contributions in NLP Papers
by: Pramanick, Aniket, et al.
Published: (2024)
by: Pramanick, Aniket, et al.
Published: (2024)
Syntax-Guided Diffusion Language Models with User-Integrated Personalization
by: Zhang, Ruqian, et al.
Published: (2025)
by: Zhang, Ruqian, et al.
Published: (2025)
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
by: Jung, Julie, et al.
Published: (2025)
by: Jung, Julie, et al.
Published: (2025)
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
by: Eckman, Stephanie, et al.
Published: (2025)
by: Eckman, Stephanie, et al.
Published: (2025)
Geological Inference from Textual Data using Word Embeddings
by: Linphrachaya, Nanmanas, et al.
Published: (2025)
by: Linphrachaya, Nanmanas, et al.
Published: (2025)
Transforming Sensitive Documents into Quantitative Data: An AI-Based Preprocessing Toolchain for Structured and Privacy-Conscious Analysis
by: Ledberg, Anders, et al.
Published: (2025)
by: Ledberg, Anders, et al.
Published: (2025)
Maximizing Signal in Human-Model Preference Alignment
by: Kraus, Kelsey, et al.
Published: (2025)
by: Kraus, Kelsey, et al.
Published: (2025)
Differential contributions of machine learning and statistical analysis to language and cognitive sciences
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
Exploring Intra and Inter-language Consistency in Embeddings with ICA
by: Li, Rongzhi, et al.
Published: (2024)
by: Li, Rongzhi, et al.
Published: (2024)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
by: Li, Yanran
Published: (2026)
by: Li, Yanran
Published: (2026)
Isolated Causal Effects of Natural Language
by: Lin, Victoria, et al.
Published: (2024)
by: Lin, Victoria, et al.
Published: (2024)
An Embedded Diachronic Sense Change Model with a Case Study from Ancient Greek
by: Zafar, Schyan, et al.
Published: (2023)
by: Zafar, Schyan, et al.
Published: (2023)
Causal Graph Discovery with Retrieval-Augmented Generation based Large Language Models
by: Zhang, Yuzhe, et al.
Published: (2024)
by: Zhang, Yuzhe, et al.
Published: (2024)
Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
Documents Are People and Words Are Items: A Psychometric Approach to Textual Data with Contextual Embeddings
by: Chen, Jinsong
Published: (2025)
by: Chen, Jinsong
Published: (2025)
Learning Dynamic Representations and Policies from Multimodal Clinical Time-Series with Informative Missingness
by: Liang, Zihan, et al.
Published: (2026)
by: Liang, Zihan, et al.
Published: (2026)
Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content
by: Bednarczyk, Lydie, et al.
Published: (2026)
by: Bednarczyk, Lydie, et al.
Published: (2026)
LEGOBench: Scientific Leaderboard Generation Benchmark
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingness
by: Liang, Zihan, et al.
Published: (2025)
by: Liang, Zihan, et al.
Published: (2025)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
by: Egami, Naoki, et al.
Published: (2023)
by: Egami, Naoki, et al.
Published: (2023)
The Deep Latent Position Topic Model for Clustering and Representation of Networks with Textual Edges
by: Boutin, Rémi, et al.
Published: (2023)
by: Boutin, Rémi, et al.
Published: (2023)
League: Leaderboard Generation on Demand
by: Wu, Jian, et al.
Published: (2025)
by: Wu, Jian, et al.
Published: (2025)
Constructing the Truth: Text Mining and Linguistic Networks in Public Hearings of Case 03 of the Special Jurisdiction for Peace (JEP)
by: Sosa, Juan, et al.
Published: (2025)
by: Sosa, Juan, et al.
Published: (2025)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
by: Hobelsberger, Christian, et al.
Published: (2025)
by: Hobelsberger, Christian, et al.
Published: (2025)
Leveraging text data for causal inference using electronic health records
by: Mozer, Reagan, et al.
Published: (2023)
by: Mozer, Reagan, et al.
Published: (2023)
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
by: Maya, Cristian Espinal
Published: (2026)
by: Maya, Cristian Espinal
Published: (2026)
Causal Inference on Outcomes Learned from Text
by: Modarressi, Iman, et al.
Published: (2025)
by: Modarressi, Iman, et al.
Published: (2025)
Automatic Debiased Machine Learning for Covariate Shifts
by: Chernozhukov, Victor, et al.
Published: (2023)
by: Chernozhukov, Victor, et al.
Published: (2023)
Similar Items
-
MetaLead: A Comprehensive Human-Curated Leaderboard Dataset for Transparent Reporting of Machine Learning Experiments
by: Timmer, Roelien C., et al.
Published: (2026) -
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025) -
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
by: Timmer, Roelien C., et al.
Published: (2024) -
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
by: Şahinuç, Furkan, et al.
Published: (2024) -
Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
by: Lyu, Y., et al.
Published: (2025)