LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Eunsu, Suk, Juyoung, Kim, Seungone, Muennighoff, Niklas, Kim, Dongkwan, Oh, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
by: Shi, Zhichao, et al.
Published: (2025)
by: Shi, Zhichao, et al.
Published: (2025)
Generalizing Weisfeiler-Lehman Kernels to Subgraphs
by: Kim, Dongkwan, et al.
Published: (2024)
by: Kim, Dongkwan, et al.
Published: (2024)
Translating Subgraphs to Nodes Makes Simple GNNs Strong and Efficient for Subgraph Representation Learning
by: Kim, Dongkwan, et al.
Published: (2022)
by: Kim, Dongkwan, et al.
Published: (2022)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
by: Ye, Seonghyeon, et al.
Published: (2023)
by: Ye, Seonghyeon, et al.
Published: (2023)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
by: Kim, Jiseon, et al.
Published: (2025)
by: Kim, Jiseon, et al.
Published: (2025)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
M-Prometheus: A Suite of Open Multilingual LLM Judges
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Better Instruction-Following Through Minimum Bayes Risk
by: Wu, Ian, et al.
Published: (2024)
by: Wu, Ian, et al.
Published: (2024)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
by: Kabir, Daeen, et al.
Published: (2025)
by: Kabir, Daeen, et al.
Published: (2025)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
by: Yoa, Seungdong, et al.
Published: (2026)
by: Yoa, Seungdong, et al.
Published: (2026)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
by: Kim, Jun Seo, et al.
Published: (2025)
by: Kim, Jun Seo, et al.
Published: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Reasoning Models Better Express Their Confidence
by: Yoon, Dongkeun, et al.
Published: (2025)
by: Yoon, Dongkeun, et al.
Published: (2025)
CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections
by: Kim, Keuntae, et al.
Published: (2025)
by: Kim, Keuntae, et al.
Published: (2025)
Beyond Static Summarization: Proactive Memory Extraction for LLM Agents
by: Yang, Chengyuan, et al.
Published: (2026)
by: Yang, Chengyuan, et al.
Published: (2026)
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
by: Lee, Yanggyu, et al.
Published: (2024)
by: Lee, Yanggyu, et al.
Published: (2024)
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
by: Kim, Hyuntak, et al.
Published: (2025)
by: Kim, Hyuntak, et al.
Published: (2025)
Trillion 7B Technical Report
by: Han, Sungjun, et al.
Published: (2025)
by: Han, Sungjun, et al.
Published: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
by: Hong, Seokhee, et al.
Published: (2025)
by: Hong, Seokhee, et al.
Published: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
by: Oh, Jihwan, et al.
Published: (2024)
by: Oh, Jihwan, et al.
Published: (2024)
Flow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through Options
by: Nair, Lakshmi, et al.
Published: (2025)
by: Nair, Lakshmi, et al.
Published: (2025)
Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation
by: Kim, Ireh, et al.
Published: (2026)
by: Kim, Ireh, et al.
Published: (2026)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
GuideLLM: Exploring LLM-Guided Conversation with Applications in Autobiography Interviewing
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
Aligning Large Language Models for Enhancing Psychiatric Interviews Through Symptom Delineation and Summarization: Pilot Study
by: So, Jae-hee, et al.
Published: (2024)
by: So, Jae-hee, et al.
Published: (2024)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
by: Park, Junyeong, et al.
Published: (2025)
by: Park, Junyeong, et al.
Published: (2025)
Similar Items
-
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025) -
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024) -
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025) -
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025) -
Uncovering Factor Level Preferences to Improve Human-Model Alignment
by: Oh, Juhyun, et al.
Published: (2024)