Economic Evaluation of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zellinger, Michael J., Thomson, Matt |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficiently Deploying LLMs with Controlled Risk
by: Zellinger, Michael J., et al.
Published: (2024)
by: Zellinger, Michael J., et al.
Published: (2024)
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Rational Tuning of LLM Cascades via Probabilistic Modeling
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Cost-Saving LLM Cascades with Early Abstention
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Herd: Using multiple, smaller LLMs to match the performances of proprietary, large LLMs via an intelligent composer
by: Hari, Surya Narayanan, et al.
Published: (2023)
by: Hari, Surya Narayanan, et al.
Published: (2023)
Learning with Noisy Labels by Adaptive Gradient-Based Outlier Removal
by: Sedova, Anastasiia, et al.
Published: (2023)
by: Sedova, Anastasiia, et al.
Published: (2023)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024)
by: Zellinger, Lena, et al.
Published: (2024)
Leveraging Open-Source Large Language Models for encoding Social Determinants of Health using an Intelligent Router
by: Goel, Akul, et al.
Published: (2024)
by: Goel, Akul, et al.
Published: (2024)
Prompt Baking
by: Bhargava, Aman, et al.
Published: (2024)
by: Bhargava, Aman, et al.
Published: (2024)
LLMs Can Assist with Proposal Selection at Large User Facilities
by: Ding, Lijie, et al.
Published: (2025)
by: Ding, Lijie, et al.
Published: (2025)
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
by: Cullen, Carissa, et al.
Published: (2026)
by: Cullen, Carissa, et al.
Published: (2026)
Evaluating Deduplication Techniques for Economic Research Paper Titles with a Focus on Semantic Similarity using NLP and LLMs
by: You, Doohee, et al.
Published: (2024)
by: You, Doohee, et al.
Published: (2024)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
by: Wachter, Jasmin, et al.
Published: (2025)
by: Wachter, Jasmin, et al.
Published: (2025)
MALLES: A Multi-agent LLMs-based Economic Sandbox with Consumer Preference Alignment
by: Wu, Yusen, et al.
Published: (2026)
by: Wu, Yusen, et al.
Published: (2026)
What's the Magic Word? A Control Theory of LLM Prompting
by: Bhargava, Aman, et al.
Published: (2023)
by: Bhargava, Aman, et al.
Published: (2023)
Evaluation of LLMs for mathematical problem solving
by: Wang, Ruonan, et al.
Published: (2025)
by: Wang, Ruonan, et al.
Published: (2025)
Graph Drawing for LLMs: An Empirical Evaluation
by: Didimo, Walter, et al.
Published: (2025)
by: Didimo, Walter, et al.
Published: (2025)
Evaluating Developmental Cognition Capabilities of LLMs
by: Xiao, Xiao, et al.
Published: (2026)
by: Xiao, Xiao, et al.
Published: (2026)
Reinforcing privacy reasoning in LLMs via normative simulacra from fiction
by: Franchi, Matt, et al.
Published: (2026)
by: Franchi, Matt, et al.
Published: (2026)
Where the Really Hard Quadratic Assignment Problems Are: the QAP-SAT instances
by: Verel, Sébastien, et al.
Published: (2024)
by: Verel, Sébastien, et al.
Published: (2024)
The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems
by: Belz, Anya, et al.
Published: (2025)
by: Belz, Anya, et al.
Published: (2025)
Fine-Tuning LLMs to Generate Economical and Reliable Actions for the Power Grid
by: Chehade, Mohamad, et al.
Published: (2026)
by: Chehade, Mohamad, et al.
Published: (2026)
SEAL: Suite for Evaluating API-use of LLMs
by: Kim, Woojeong, et al.
Published: (2024)
by: Kim, Woojeong, et al.
Published: (2024)
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
by: Sun, Zhaoyue, et al.
Published: (2026)
by: Sun, Zhaoyue, et al.
Published: (2026)
Navigating Pitfalls: Evaluating LLMs in Machine Learning Programming Education
by: Kumar, Smitha, et al.
Published: (2025)
by: Kumar, Smitha, et al.
Published: (2025)
SymbolicAI: A framework for logic-based approaches combining generative models and solvers
by: Dinu, Marius-Constantin, et al.
Published: (2024)
by: Dinu, Marius-Constantin, et al.
Published: (2024)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023)
by: Raunak, Vikas, et al.
Published: (2023)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
Evaluating LLMs for Visualization Tasks
by: Khan, Saadiq Rauf, et al.
Published: (2025)
by: Khan, Saadiq Rauf, et al.
Published: (2025)
Cost-of-Pass: An Economic Framework for Evaluating Language Models
by: Erol, Mehmet Hamza, et al.
Published: (2025)
by: Erol, Mehmet Hamza, et al.
Published: (2025)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
by: Li, Hanyu, et al.
Published: (2025)
by: Li, Hanyu, et al.
Published: (2025)
OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data
by: Renda, Alana, et al.
Published: (2025)
by: Renda, Alana, et al.
Published: (2025)
R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
by: Jain, Raj, et al.
Published: (2025)
by: Jain, Raj, et al.
Published: (2025)
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
by: Yueh-Han, Chen, et al.
Published: (2025)
by: Yueh-Han, Chen, et al.
Published: (2025)
Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
by: Galatolo, Alessio, et al.
Published: (2025)
by: Galatolo, Alessio, et al.
Published: (2025)
Evaluating Prompting and Execution-Based Methods for Deterministic Computation in LLMs
by: Yu, Hongkun
Published: (2026)
by: Yu, Hongkun
Published: (2026)
Evaluating Explanations Through LLMs: Beyond Traditional User Studies
by: De Bona, Francesco Bombassei, et al.
Published: (2024)
by: De Bona, Francesco Bombassei, et al.
Published: (2024)
Evaluating LLMs for Answering Student Questions in Introductory Programming Courses
by: Van Mullem, Thomas, et al.
Published: (2026)
by: Van Mullem, Thomas, et al.
Published: (2026)
Similar Items
-
Efficiently Deploying LLMs with Controlled Risk
by: Zellinger, Michael J., et al.
Published: (2024) -
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
by: Zellinger, Michael J., et al.
Published: (2025) -
Rational Tuning of LLM Cascades via Probabilistic Modeling
by: Zellinger, Michael J., et al.
Published: (2025) -
Cost-Saving LLM Cascades with Early Abstention
by: Zellinger, Michael J., et al.
Published: (2025) -
Herd: Using multiple, smaller LLMs to match the performances of proprietary, large LLMs via an intelligent composer
by: Hari, Surya Narayanan, et al.
Published: (2023)