Evalet: Evaluating Large Language Models through Functional Fragmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Tae Soo, Lee, Heechan, Lee, Yoonjoo, Seering, Joseph, Kim, Juho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
"When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction
von: Son, Kihoon, et al.
Veröffentlicht: (2026)
von: Son, Kihoon, et al.
Veröffentlicht: (2026)
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale Inference
von: Son, Kihoon, et al.
Veröffentlicht: (2025)
von: Son, Kihoon, et al.
Veröffentlicht: (2025)
Inertia in Moral and Value Judgments of Large Language Models
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
VIVID: Human-AI Collaborative Authoring of Vicarious Dialogues from Lecture Videos
von: Choi, Seulgi, et al.
Veröffentlicht: (2024)
von: Choi, Seulgi, et al.
Veröffentlicht: (2024)
Chain of Empathy: Enhancing Empathetic Response of Large Language Models Based on Psychotherapy Models
von: Lee, Yoon Kyung, et al.
Veröffentlicht: (2023)
von: Lee, Yoon Kyung, et al.
Veröffentlicht: (2023)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
von: Yoon, Sion, et al.
Veröffentlicht: (2024)
von: Yoon, Sion, et al.
Veröffentlicht: (2024)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
The Adoption and Efficacy of Large Language Models: Evidence From Consumer Complaints in the Financial Industry
von: Shin, Minkyu, et al.
Veröffentlicht: (2023)
von: Shin, Minkyu, et al.
Veröffentlicht: (2023)
Prototyping Digital Social Spaces through Metaphor-Driven Design: Translating Spatial Concepts into an Interactive Social Simulation
von: Hong, Yoojin, et al.
Veröffentlicht: (2025)
von: Hong, Yoojin, et al.
Veröffentlicht: (2025)
An Evaluation of Estimative Uncertainty in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
"Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students using Large Language Models
von: Bernstein, Seth, et al.
Veröffentlicht: (2024)
von: Bernstein, Seth, et al.
Veröffentlicht: (2024)
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models
von: Lee, Yukyung, et al.
Veröffentlicht: (2024)
von: Lee, Yukyung, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in Analysing Classroom Dialogue
von: Long, Yun, et al.
Veröffentlicht: (2024)
von: Long, Yun, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
von: Kim, Yoonsu, et al.
Veröffentlicht: (2025)
von: Kim, Yoonsu, et al.
Veröffentlicht: (2025)
Comparing Code Explanations Created by Students and Large Language Models
von: Leinonen, Juho, et al.
Veröffentlicht: (2023)
von: Leinonen, Juho, et al.
Veröffentlicht: (2023)
Iffy-Or-Not: Extending the Web to Support the Critical Evaluation of Fallacious Texts
von: Lim, Gionnieve, et al.
Veröffentlicht: (2025)
von: Lim, Gionnieve, et al.
Veröffentlicht: (2025)
Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf Game
von: Du, Silin, et al.
Veröffentlicht: (2024)
von: Du, Silin, et al.
Veröffentlicht: (2024)
The Effectiveness of Style Vectors for Steering Large Language Models: A Human Evaluation
von: Diallo, Diaoulé, et al.
Veröffentlicht: (2026)
von: Diallo, Diaoulé, et al.
Veröffentlicht: (2026)
Analyzing Cognitive Differences Among Large Language Models through the Lens of Social Worldview
von: Li, Jiatao, et al.
Veröffentlicht: (2025)
von: Li, Jiatao, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Hybrid Workplace Decision Support
von: Kim, Yujin, et al.
Veröffentlicht: (2024)
von: Kim, Yujin, et al.
Veröffentlicht: (2024)
Using Large Language Models to Enhance Programming Error Messages
von: Leinonen, Juho, et al.
Veröffentlicht: (2022)
von: Leinonen, Juho, et al.
Veröffentlicht: (2022)
Human Evaluation of Procedural Knowledge Graph Extraction from Text with Large Language Models
von: Carriero, Valentina Anita, et al.
Veröffentlicht: (2024)
von: Carriero, Valentina Anita, et al.
Veröffentlicht: (2024)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
von: Lee, Minhwa, et al.
Veröffentlicht: (2024)
von: Lee, Minhwa, et al.
Veröffentlicht: (2024)
Moving Beyond Review: Applying Language Models to Planning and Translation in Reflection
von: Neshaei, Seyed Parsa, et al.
Veröffentlicht: (2026)
von: Neshaei, Seyed Parsa, et al.
Veröffentlicht: (2026)
From Divergence to Consensus: Evaluating the Role of Large Language Models in Facilitating Agreement through Adaptive Strategies
von: Triantafyllopoulos, Loukas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Loukas, et al.
Veröffentlicht: (2025)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
von: Sun, Chongyan, et al.
Veröffentlicht: (2024)
von: Sun, Chongyan, et al.
Veröffentlicht: (2024)
HealthGenie: Empowering Users with Healthy Dietary Guidance through Knowledge Graph and Large Language Models
von: Gao, Fan, et al.
Veröffentlicht: (2025)
von: Gao, Fan, et al.
Veröffentlicht: (2025)
LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
von: Liu, Minqian, et al.
Veröffentlicht: (2025)
von: Liu, Minqian, et al.
Veröffentlicht: (2025)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
Collaborative Evaluation of Deepfake Text with Deliberation-Enhancing Dialogue Systems
von: Lee, Jooyoung, et al.
Veröffentlicht: (2025)
von: Lee, Jooyoung, et al.
Veröffentlicht: (2025)
Less Talk, More Trust: Understanding Players' In-game Assessment of Communication Processes in League of Legends
von: Lee, Juhoon, et al.
Veröffentlicht: (2025)
von: Lee, Juhoon, et al.
Veröffentlicht: (2025)
Epistemic Integrity in Large Language Models
von: Ghafouri, Bijean, et al.
Veröffentlicht: (2024)
von: Ghafouri, Bijean, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models' Ability Using a Psychiatric Screening Tool Based on Metaphor and Sarcasm Scenarios
von: Yakura, Hiromu
Veröffentlicht: (2023)
von: Yakura, Hiromu
Veröffentlicht: (2023)
Ähnliche Einträge
-
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023) -
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025) -
DiscoverLLM: From Executing Intents to Discovering Them
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026) -
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024) -
"When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction
von: Son, Kihoon, et al.
Veröffentlicht: (2026)