Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
Fuente:
arXiv
Guardado en:
| Autores principales: | Murugadoss, Bhuvanashree, Poelitz, Christian, Drosos, Ian, Le, Vu, McKenna, Nick, Negreanu, Carina Suzana, Parnin, Chris, Sarkar, Advait |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Synthetic Clarification and Correction Dialogues about Data-Centric Tasks -- A Teacher-Student Approach
por: Poelitz, Christian, et al.
Publicado: (2025)
por: Poelitz, Christian, et al.
Publicado: (2025)
When Copilot Becomes Autopilot: Generative AI's Critical Risk to Knowledge Work and a Critical Solution
por: Sarkar, Advait, et al.
Publicado: (2024)
por: Sarkar, Advait, et al.
Publicado: (2024)
Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
por: Kazemitabaar, Majeed, et al.
Publicado: (2024)
por: Kazemitabaar, Majeed, et al.
Publicado: (2024)
"It's like a rubber duck that talks back": Understanding Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study
por: Drosos, Ian, et al.
Publicado: (2024)
por: Drosos, Ian, et al.
Publicado: (2024)
Vibe coding: programming through conversation with artificial intelligence
por: Sarkar, Advait, et al.
Publicado: (2025)
por: Sarkar, Advait, et al.
Publicado: (2025)
Solving Data-centric Tasks using Large Language Models
por: Barke, Shraddha, et al.
Publicado: (2024)
por: Barke, Shraddha, et al.
Publicado: (2024)
Dynamic Prompt Middleware: Contextual Prompt Refinement Controls for Comprehension Tasks
por: Drosos, Ian, et al.
Publicado: (2024)
por: Drosos, Ian, et al.
Publicado: (2024)
From Binary Groundedness to Support Relations: Towards a Reader-Centred Taxonomy for Comprehension of AI Output
por: Sarkar, Advait, et al.
Publicado: (2026)
por: Sarkar, Advait, et al.
Publicado: (2026)
"My toxic trait is thinking I'll remember this": gaps in the learner experience of video tutorials for feature-rich software
por: Drosos, Ian, et al.
Publicado: (2024)
por: Drosos, Ian, et al.
Publicado: (2024)
An Experimental Comparison of Cognitive Forcing Functions for Execution Plans in AI-Assisted Writing: Effects On Trust, Overreliance, and Perceived Critical Thinking
por: Ghosh, Ahana, et al.
Publicado: (2026)
por: Ghosh, Ahana, et al.
Publicado: (2026)
"It makes you think": Provocations Help Restore Critical Thinking to AI-Assisted Knowledge Work
por: Drosos, Ian, et al.
Publicado: (2025)
por: Drosos, Ian, et al.
Publicado: (2025)
Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages
por: McKenna, Nick, et al.
Publicado: (2025)
por: McKenna, Nick, et al.
Publicado: (2025)
How Is Scale Incorporated Into the Economic Evaluation of Interventions to Prevent Obesity or to Improve Obesity‐Related Risk Factors: A Systematic Scoping Review
por: Carina Dalton, et al.
Publicado: (2025)
por: Carina Dalton, et al.
Publicado: (2025)
Scaling up the Banded Matrix Factorization Mechanism for Differentially Private ML
por: McKenna, Ryan
Publicado: (2024)
por: McKenna, Ryan
Publicado: (2024)
Making Sense of Sensemaking: Designing Authentic K-12 STEM Learning Experiences
por: T. J. McKenna
Publicado: (2025)
por: T. J. McKenna
Publicado: (2025)
Libraries and the Internet. ERIC Digest.
por: McKenna, Mary
Publicado: (1994)
por: McKenna, Mary
Publicado: (1994)
So Many Students, so Little Time: Practical Student Worker Training in an Academic Library
por: McKenna, Julia
Publicado: (2020)
por: McKenna, Julia
Publicado: (2020)
HyLife: The Hybrid Library of the Future.
por: McKenna, Brian
Publicado: (1999)
por: McKenna, Brian
Publicado: (1999)
ANTHROPOLOGY MUST EMBRACE JOURNALISM. PUBLIC PEDAGOGY IS DISCIPLINE'S CHALLENGE
por: Brian McKenna
Publicado: (2010)
por: Brian McKenna
Publicado: (2010)
Beyond the Comfort Zone: Emerging Solutions to Overcome Challenges in Integrating LLMs into Software Products
por: Nahar, Nadia, et al.
Publicado: (2024)
por: Nahar, Nadia, et al.
Publicado: (2024)
AI Should Challenge, Not Obey
por: Sarkar, Advait
Publicado: (2024)
por: Sarkar, Advait
Publicado: (2024)
Searching for European Alternatives: Digital Sovereignty, Digital Patriotism, and the Emerging Geopolitics of Software Adoption
por: Sarkar, Advait
Publicado: (2026)
por: Sarkar, Advait
Publicado: (2026)
Filter Babel: The Challenge of Synthetic Media to Authenticity and Common Ground in AI-Mediated Communication
por: Sarkar, Advait
Publicado: (2026)
por: Sarkar, Advait
Publicado: (2026)
Large Language Models Cannot Explain Themselves
por: Sarkar, Advait
Publicado: (2024)
por: Sarkar, Advait
Publicado: (2024)
Intention Is All You Need
por: Sarkar, Advait
Publicado: (2024)
por: Sarkar, Advait
Publicado: (2024)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
por: Enguehard, Joseph, et al.
Publicado: (2025)
por: Enguehard, Joseph, et al.
Publicado: (2025)
The Stability Trap: Evaluating the Reliability of LLM-Based Instruction Adherence Auditing
por: Shergadwala, Murtuza N.
Publicado: (2026)
por: Shergadwala, Murtuza N.
Publicado: (2026)
Using Micros to Find Fiction: Issues and Answers.
por: McKenna, Michael C.
Publicado: (1987)
por: McKenna, Michael C.
Publicado: (1987)
The Humanities Flourish in Rural Communities.
por: McKenna, Paul G.
Publicado: (1983)
por: McKenna, Paul G.
Publicado: (1983)
Bonaventure’s ‘Journey of the Soul into God’: Context and Commentary. By RandallSmith. Cambridge: Cambridge University Press, 2025. Pp. 512. £120.00
por: Thomas J. McKenna
Publicado: (2025)
por: Thomas J. McKenna
Publicado: (2025)
BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios
por: Tilak, Advait, et al.
Publicado: (2026)
por: Tilak, Advait, et al.
Publicado: (2026)
Beyond Diagnosis: Evaluating Multimodal LLMs for Pathology Localization in Chest Radiographs
por: Gosai, Advait, et al.
Publicado: (2025)
por: Gosai, Advait, et al.
Publicado: (2025)
Evaluating and Enhancing Trustworthiness of LLMs in Perception Tasks
por: Dona, Malsha Ashani Mahawatta, et al.
Publicado: (2024)
por: Dona, Malsha Ashani Mahawatta, et al.
Publicado: (2024)
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
por: Yan, Litao, et al.
Publicado: (2025)
por: Yan, Litao, et al.
Publicado: (2025)
Evaluating LLMs for Visualization Tasks
por: Khan, Saadiq Rauf, et al.
Publicado: (2025)
por: Khan, Saadiq Rauf, et al.
Publicado: (2025)
Accompaniment Prompt Adherence: A Measure for Evaluating Music Accompaniment Systems
por: Grachten, Maarten, et al.
Publicado: (2025)
por: Grachten, Maarten, et al.
Publicado: (2025)
Safety Evaluation of Repeated Application of Polymeric Microarray Patches in Miniature Pigs
por: Qonita Kurnia Anjani, et al.
Publicado: (2025)
por: Qonita Kurnia Anjani, et al.
Publicado: (2025)
A moment-based approach to the injective norm of random tensors
por: Dartois, Stephane, et al.
Publicado: (2026)
por: Dartois, Stephane, et al.
Publicado: (2026)
Understanding Higher Education: Alternative Perspectives
por: Boughey, Chrissie, et al.
Publicado: (2021)
por: Boughey, Chrissie, et al.
Publicado: (2021)
The duty to listen
por: Hrishikesh Joshi, et al.
Publicado: (2024)
por: Hrishikesh Joshi, et al.
Publicado: (2024)
Ejemplares similares
-
Synthetic Clarification and Correction Dialogues about Data-Centric Tasks -- A Teacher-Student Approach
por: Poelitz, Christian, et al.
Publicado: (2025) -
When Copilot Becomes Autopilot: Generative AI's Critical Risk to Knowledge Work and a Critical Solution
por: Sarkar, Advait, et al.
Publicado: (2024) -
Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
por: Kazemitabaar, Majeed, et al.
Publicado: (2024) -
"It's like a rubber duck that talks back": Understanding Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study
por: Drosos, Ian, et al.
Publicado: (2024) -
Vibe coding: programming through conversation with artificial intelligence
por: Sarkar, Advait, et al.
Publicado: (2025)