Can Language Models Evaluate Human Written Text? Case Study on Korean Student Writing for Education
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Seungyoon, Kim, Seungone |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
von: Kim, Seoyeon, et al.
Veröffentlicht: (2024)
von: Kim, Seoyeon, et al.
Veröffentlicht: (2024)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023)
von: Son, Guijin, et al.
Veröffentlicht: (2023)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2024)
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2024)
LLM-as-a-tutor in EFL Writing Education: Focusing on Evaluation of Student-LLM Interaction
von: Han, Jieun, et al.
Veröffentlicht: (2023)
von: Han, Jieun, et al.
Veröffentlicht: (2023)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2023)
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2023)
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected?
von: Zhang, Qihui, et al.
Veröffentlicht: (2024)
von: Zhang, Qihui, et al.
Veröffentlicht: (2024)
Evaluating Multimodal Generative AI with Korean Educational Standards
von: Park, Sanghee, et al.
Veröffentlicht: (2025)
von: Park, Sanghee, et al.
Veröffentlicht: (2025)
Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
von: Sasagawa, Keito, et al.
Veröffentlicht: (2025)
von: Sasagawa, Keito, et al.
Veröffentlicht: (2025)
Optimizing Language Augmentation for Multilingual Large Language Models: A Case Study on Korean
von: Choi, ChangSu, et al.
Veröffentlicht: (2024)
von: Choi, ChangSu, et al.
Veröffentlicht: (2024)
Measuring Sycophancy of Language Models in Multi-turn Dialogues
von: Hong, Jiseung, et al.
Veröffentlicht: (2025)
von: Hong, Jiseung, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
Evaluating Language Models as Synthetic Data Generators
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
GECKO: Generative Language Model for English, Code and Korean
von: Oh, Sungwoo, et al.
Veröffentlicht: (2024)
von: Oh, Sungwoo, et al.
Veröffentlicht: (2024)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
von: Ju, Jeongho, et al.
Veröffentlicht: (2024)
von: Ju, Jeongho, et al.
Veröffentlicht: (2024)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
UKTA: Unified Korean Text Analyzer
von: Ahn, Seokho, et al.
Veröffentlicht: (2025)
von: Ahn, Seokho, et al.
Veröffentlicht: (2025)
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
von: Lim, Junghwan, et al.
Veröffentlicht: (2025)
von: Lim, Junghwan, et al.
Veröffentlicht: (2025)
Chain-of-MetaWriting: Linguistic and Textual Analysis of How Small Language Models Write Young Students Texts
von: Buhnila, Ioana, et al.
Veröffentlicht: (2024)
von: Buhnila, Ioana, et al.
Veröffentlicht: (2024)
Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
Does Incomplete Syntax Influence Korean Language Model? Focusing on Word Order and Case Markers
von: Kim, Jong Myoung, et al.
Veröffentlicht: (2024)
von: Kim, Jong Myoung, et al.
Veröffentlicht: (2024)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2023)
von: Kim, Seungone, et al.
Veröffentlicht: (2023)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
ChEDDAR: Student-ChatGPT Dialogue in EFL Writing Education
von: Han, Jieun, et al.
Veröffentlicht: (2023)
von: Han, Jieun, et al.
Veröffentlicht: (2023)
Can Large Language Models Automatically Score Proficiency of Written Essays?
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
Theme-Explanation Structure for Table Summarization using Large Language Models: A Case Study on Korean Tabular Data
von: Kwack, TaeYoon, et al.
Veröffentlicht: (2025)
von: Kwack, TaeYoon, et al.
Veröffentlicht: (2025)
Linguistically Informed Graph Model and Semantic Contrastive Learning for Korean Short Text Classification
von: Yoo, JaeGeon, et al.
Veröffentlicht: (2026)
von: Yoo, JaeGeon, et al.
Veröffentlicht: (2026)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
von: Fang, Hao, et al.
Veröffentlicht: (2025)
von: Fang, Hao, et al.
Veröffentlicht: (2025)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
von: Lee, Minhwa, et al.
Veröffentlicht: (2024)
von: Lee, Minhwa, et al.
Veröffentlicht: (2024)
Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
Optimizing Korean-Centric LLMs via Token Pruning
von: Kim, Hoyeol, et al.
Veröffentlicht: (2026)
von: Kim, Hoyeol, et al.
Veröffentlicht: (2026)
Low-Resource NMT: A Case Study on the Written and Spoken Languages in Hong Kong
von: Mak, Hei Yi, et al.
Veröffentlicht: (2025)
von: Mak, Hei Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
von: Kim, Seoyeon, et al.
Veröffentlicht: (2024) -
KMMLU: Measuring Massive Multitask Language Understanding in Korean
von: Son, Guijin, et al.
Veröffentlicht: (2024) -
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
von: Son, Guijin, et al.
Veröffentlicht: (2024) -
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
von: Lee, Seongyun, et al.
Veröffentlicht: (2024) -
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023)