A Practical Guide for Evaluating LLMs and LLM-Reliant Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Rudd, Ethan M., Andrews, Christopher, Tully, Philip |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On Distributional Reinforcement Learning in Chaotic Dynamical Systems
por: Rudd-Jones, James, et al.
Publicado: (2026)
por: Rudd-Jones, James, et al.
Publicado: (2026)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
por: Liang, Zhenwen, et al.
Publicado: (2025)
por: Liang, Zhenwen, et al.
Publicado: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
From Theory to Practice: Implementing and Evaluating e-Fold Cross-Validation
por: Mahlich, Christopher, et al.
Publicado: (2024)
por: Mahlich, Christopher, et al.
Publicado: (2024)
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
por: Gangavarapu, Tushaar, et al.
Publicado: (2026)
por: Gangavarapu, Tushaar, et al.
Publicado: (2026)
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
por: Vishwakarma, Harsh, et al.
Publicado: (2025)
por: Vishwakarma, Harsh, et al.
Publicado: (2025)
Fine-tuning Timeseries Predictors Using Reinforcement Learning
por: Cazaux, Hugo, et al.
Publicado: (2026)
por: Cazaux, Hugo, et al.
Publicado: (2026)
Are LLMs The Way Forward? A Case Study on LLM-Guided Reinforcement Learning for Decentralized Autonomous Driving
por: Anvar, Timur, et al.
Publicado: (2025)
por: Anvar, Timur, et al.
Publicado: (2025)
VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
por: Smyth, Ethan, et al.
Publicado: (2025)
por: Smyth, Ethan, et al.
Publicado: (2025)
Evidence for Limited Metacognition in LLMs
por: Ackerman, Christopher
Publicado: (2025)
por: Ackerman, Christopher
Publicado: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
por: Li, Xiaochuan, et al.
Publicado: (2025)
por: Li, Xiaochuan, et al.
Publicado: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
por: Knoop, Jonathan, et al.
Publicado: (2026)
por: Knoop, Jonathan, et al.
Publicado: (2026)
One-shot Federated Learning Methods: A Practical Guide
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
por: Tang, Bohan, et al.
Publicado: (2025)
por: Tang, Bohan, et al.
Publicado: (2025)
Learning in Context, Guided by Choice: A Reward-Free Paradigm for Reinforcement Learning with Transformers
por: Dong, Juncheng, et al.
Publicado: (2026)
por: Dong, Juncheng, et al.
Publicado: (2026)
A Comprehensive Guide to Explainable AI: From Classical Models to LLMs
por: Hsieh, Weiche, et al.
Publicado: (2024)
por: Hsieh, Weiche, et al.
Publicado: (2024)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
por: Eshuijs, Leon, et al.
Publicado: (2025)
por: Eshuijs, Leon, et al.
Publicado: (2025)
A Guide to Failure in Machine Learning: Reliability and Robustness from Foundations to Practice
por: Heim, Eric, et al.
Publicado: (2025)
por: Heim, Eric, et al.
Publicado: (2025)
Attacks and Defenses Against LLM Fingerprinting
por: Kurian, Kevin, et al.
Publicado: (2025)
por: Kurian, Kevin, et al.
Publicado: (2025)
Property-Guided LLM Program Synthesis for Planning
por: Pereira, André G., et al.
Publicado: (2026)
por: Pereira, André G., et al.
Publicado: (2026)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2026)
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2026)
A Justice Lens on Fairness and Ethics Courses in Computing Education: LLM-Assisted Multi-Perspective and Thematic Evaluation
por: Andrews, Kenya S., et al.
Publicado: (2025)
por: Andrews, Kenya S., et al.
Publicado: (2025)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
por: Zhang, Yuanhe, et al.
Publicado: (2025)
por: Zhang, Yuanhe, et al.
Publicado: (2025)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
por: Cinquin, Tristan, et al.
Publicado: (2025)
por: Cinquin, Tristan, et al.
Publicado: (2025)
REASONING COMPILER: LLM-Guided Optimizations for Efficient Model Serving
por: Tang, Annabelle Sujun, et al.
Publicado: (2025)
por: Tang, Annabelle Sujun, et al.
Publicado: (2025)
How Can LLM Guide RL? A Value-Based Approach
por: Zhang, Shenao, et al.
Publicado: (2024)
por: Zhang, Shenao, et al.
Publicado: (2024)
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
por: Zimmer, Max, et al.
Publicado: (2026)
por: Zimmer, Max, et al.
Publicado: (2026)
Evaluation and Benchmarking of LLM Agents: A Survey
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
Mutation-Guided LLM-based Test Generation at Meta
por: Foster, Christopher, et al.
Publicado: (2025)
por: Foster, Christopher, et al.
Publicado: (2025)
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
por: Huang, Hanxian, et al.
Publicado: (2026)
por: Huang, Hanxian, et al.
Publicado: (2026)
Policy Guided Tree Search for Enhanced LLM Reasoning
por: Li, Yang
Publicado: (2025)
por: Li, Yang
Publicado: (2025)
OKG-LLM: Aligning Ocean Knowledge Graph with Observation Data via LLMs for Global Sea Surface Temperature Prediction
por: Yang, Hanchen, et al.
Publicado: (2025)
por: Yang, Hanchen, et al.
Publicado: (2025)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
por: Rezaei, Mohammad, et al.
Publicado: (2026)
por: Rezaei, Mohammad, et al.
Publicado: (2026)
Process Reward Models for LLM Agents: Practical Framework and Directions
por: Choudhury, Sanjiban
Publicado: (2025)
por: Choudhury, Sanjiban
Publicado: (2025)
Robust Guided Diffusion for Offline Black-Box Optimization
por: Chen, Can Sam, et al.
Publicado: (2024)
por: Chen, Can Sam, et al.
Publicado: (2024)
On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2024)
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2024)
Practical Bayesian Inference for Speech SNNs: Uncertainty and Loss-Landscape Smoothing
por: Abdennadher, Yesmine, et al.
Publicado: (2026)
por: Abdennadher, Yesmine, et al.
Publicado: (2026)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
por: Andrews, Martin, et al.
Publicado: (2025)
por: Andrews, Martin, et al.
Publicado: (2025)
FHE-Agent: Automating CKKS Configuration for Practical Encrypted Inference via an LLM-Guided Agentic Framework
por: Xu, Nuo, et al.
Publicado: (2025)
por: Xu, Nuo, et al.
Publicado: (2025)
Low Variance Off-policy Evaluation with State-based Importance Sampling
por: Bossens, David M., et al.
Publicado: (2022)
por: Bossens, David M., et al.
Publicado: (2022)
Ejemplares similares
-
On Distributional Reinforcement Learning in Chaotic Dynamical Systems
por: Rudd-Jones, James, et al.
Publicado: (2026) -
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
por: Liang, Zhenwen, et al.
Publicado: (2025) -
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025) -
From Theory to Practice: Implementing and Evaluating e-Fold Cross-Validation
por: Mahlich, Christopher, et al.
Publicado: (2024) -
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
por: Gangavarapu, Tushaar, et al.
Publicado: (2026)