Forest vs Tree: The $(N, K)$ Trade-off in Reproducible ML Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Pandita, Deepak, Korn, Flip, Welty, Chris, Homan, Christopher M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
by: Pandita, Deepak, et al.
Published: (2026)
by: Pandita, Deepak, et al.
Published: (2026)
How Many Ratings per Item are Necessary for Reliable Significance Testing?
by: Homan, Christopher, et al.
Published: (2024)
by: Homan, Christopher, et al.
Published: (2024)
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
by: Sawkar, Mandira, et al.
Published: (2025)
by: Sawkar, Mandira, et al.
Published: (2025)
ProRefine: Inference-Time Prompt Refinement with Textual Feedback
by: Pandita, Deepak, et al.
Published: (2025)
by: Pandita, Deepak, et al.
Published: (2025)
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
by: Shetty, Samay U., et al.
Published: (2026)
by: Shetty, Samay U., et al.
Published: (2026)
Understanding the Quality-Diversity Trade-off in Diffusion Language Models
by: Buzzard, Zak
Published: (2025)
by: Buzzard, Zak
Published: (2025)
Text to Trust: Evaluating Fine-Tuning and LoRA Trade-offs in Language Models for Unfair Terms of Service Detection
by: Juttu, Noshitha Padma Pratyusha, et al.
Published: (2025)
by: Juttu, Noshitha Padma Pratyusha, et al.
Published: (2025)
Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales
by: Resck, Lucas E., et al.
Published: (2024)
by: Resck, Lucas E., et al.
Published: (2024)
Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models
by: Chen, Pin-Yu, et al.
Published: (2025)
by: Chen, Pin-Yu, et al.
Published: (2025)
ARTICLE: Annotator Reliability Through In-Context Learning
by: Dutta, Sujan, et al.
Published: (2024)
by: Dutta, Sujan, et al.
Published: (2024)
Exploring Accuracy-Fairness Trade-off in Large Language Models
by: Zhang, Qingquan, et al.
Published: (2024)
by: Zhang, Qingquan, et al.
Published: (2024)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Emissions and Performance Trade-off Between Small and Large Language Models
by: Garg, Anandita, et al.
Published: (2025)
by: Garg, Anandita, et al.
Published: (2025)
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
by: Rathore, Darshita, et al.
Published: (2025)
by: Rathore, Darshita, et al.
Published: (2025)
Evaluating Cross-Lingual Classification Approaches Enabling Topic Discovery for Multilingual Social Media Data
by: Uniyal, Deepak, et al.
Published: (2026)
by: Uniyal, Deepak, et al.
Published: (2026)
Downstream Trade-offs of a Family of Text Watermarks
by: Ajith, Anirudh, et al.
Published: (2023)
by: Ajith, Anirudh, et al.
Published: (2023)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025)
by: Zeng, Qiuhai, et al.
Published: (2025)
Algebraic Quantum Intelligence: A New Framework for Reproducible Machine Creativity
by: Yano, Kazuo, et al.
Published: (2026)
by: Yano, Kazuo, et al.
Published: (2026)
Kernel Banzhaf: A Fast and Robust Estimator for Banzhaf Values
by: Liu, Yurong, et al.
Published: (2024)
by: Liu, Yurong, et al.
Published: (2024)
JAF: Judge Agent Forest
by: Garg, Sahil, et al.
Published: (2026)
by: Garg, Sahil, et al.
Published: (2026)
Quantifying True Robustness: Synonymity-Weighted Similarity for Trustworthy XAI Evaluation
by: Burger, Christopher
Published: (2025)
by: Burger, Christopher
Published: (2025)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
by: Fonseca, Joao, et al.
Published: (2025)
by: Fonseca, Joao, et al.
Published: (2025)
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
by: Ben-Zion, Ziv, et al.
Published: (2025)
by: Ben-Zion, Ziv, et al.
Published: (2025)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
by: Qiu, Jiahao, et al.
Published: (2024)
by: Qiu, Jiahao, et al.
Published: (2024)
CORE: Comprehensive Ontological Relation Evaluation for Large Language Models
by: Dwivedi, Satyam, et al.
Published: (2026)
by: Dwivedi, Satyam, et al.
Published: (2026)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
by: Georganas, Evangelos, et al.
Published: (2025)
by: Georganas, Evangelos, et al.
Published: (2025)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
Exploring the Performance of ML/DL Architectures on the MNIST-1D Dataset
by: Beebe, Michael, et al.
Published: (2026)
by: Beebe, Michael, et al.
Published: (2026)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
Rethinking Scale: The Efficacy of Fine-Tuned Open-Source LLMs in Large-Scale Reproducible Social Science Research
by: Carammia, Marcello, et al.
Published: (2024)
by: Carammia, Marcello, et al.
Published: (2024)
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
by: Vu, Tu, et al.
Published: (2024)
by: Vu, Tu, et al.
Published: (2024)
$\mathbf{(N,K)}$-Puzzle: A Cost-Efficient Testbed for Benchmarking Reinforcement Learning Algorithms in Generative Language Model
by: Zhang, Yufeng, et al.
Published: (2024)
by: Zhang, Yufeng, et al.
Published: (2024)
Universal Transformers Need Memory: Depth-State Trade-offs in Adaptive Recursive Reasoning
by: Sapunov, Grigory
Published: (2026)
by: Sapunov, Grigory
Published: (2026)
AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
by: Lee, Sangjun, et al.
Published: (2025)
by: Lee, Sangjun, et al.
Published: (2025)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
by: Ivanova, Anna A., et al.
Published: (2024)
by: Ivanova, Anna A., et al.
Published: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
by: Li, Aochong Oliver, et al.
Published: (2025)
by: Li, Aochong Oliver, et al.
Published: (2025)
Similar Items
-
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
by: Pandita, Deepak, et al.
Published: (2026) -
How Many Ratings per Item are Necessary for Reliable Significance Testing?
by: Homan, Christopher, et al.
Published: (2024) -
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
by: Sawkar, Mandira, et al.
Published: (2025) -
ProRefine: Inference-Time Prompt Refinement with Textual Feedback
by: Pandita, Deepak, et al.
Published: (2025) -
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
by: Shetty, Samay U., et al.
Published: (2026)