The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kostiuk, Yevhen, Vitman, Oxana, Gagała, Łukasz, Kiulian, Artur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2025)
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2025)
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
Dialectical Behavior Therapy Approach to LLM Prompting
von: Vitman, Oxana, et al.
Veröffentlicht: (2024)
von: Vitman, Oxana, et al.
Veröffentlicht: (2024)
Towards LLM-based Autograding for Short Textual Answers
von: Schneider, Johannes, et al.
Veröffentlicht: (2023)
von: Schneider, Johannes, et al.
Veröffentlicht: (2023)
Integrated Framework for LLM Evaluation with Answer Generation
von: Lee, Sujeong, et al.
Veröffentlicht: (2025)
von: Lee, Sujeong, et al.
Veröffentlicht: (2025)
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
Directed Graph-alignment Approach for Identification of Gaps in Short Answers
von: Sahu, Archana, et al.
Veröffentlicht: (2025)
von: Sahu, Archana, et al.
Veröffentlicht: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
von: Cencerrado, Iván Vicente Moreno, et al.
Veröffentlicht: (2025)
von: Cencerrado, Iván Vicente Moreno, et al.
Veröffentlicht: (2025)
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2026)
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2026)
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
von: Li, Dawei, et al.
Veröffentlicht: (2024)
von: Li, Dawei, et al.
Veröffentlicht: (2024)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
von: Chen, Luyu, et al.
Veröffentlicht: (2025)
von: Chen, Luyu, et al.
Veröffentlicht: (2025)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
von: Islam, Saad Obaid ul, et al.
Veröffentlicht: (2025)
von: Islam, Saad Obaid ul, et al.
Veröffentlicht: (2025)
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
von: Oh, Jungsuk, et al.
Veröffentlicht: (2025)
von: Oh, Jungsuk, et al.
Veröffentlicht: (2025)
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
von: Ye, Bingyang, et al.
Veröffentlicht: (2026)
von: Ye, Bingyang, et al.
Veröffentlicht: (2026)
Evaluating and Optimizing Educational Content with Large Language Model Judgments
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
von: He-Yueya, Joy, et al.
Veröffentlicht: (2024)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
Beyond Scores: A Modular RAG-Based System for Automatic Short Answer Scoring with Feedback
von: Fateen, Menna, et al.
Veröffentlicht: (2024)
von: Fateen, Menna, et al.
Veröffentlicht: (2024)
ASAG2024: A Combined Benchmark for Short Answer Grading
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024)
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
von: Aggarwal, Dishank, et al.
Veröffentlicht: (2024)
von: Aggarwal, Dishank, et al.
Veröffentlicht: (2024)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2025)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2025)
The Effect of Document Summarization on LLM-Based Relevance Judgments
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2025)
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2025)
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
von: Schleifer, Abigail Victoria Gurin, et al.
Veröffentlicht: (2026)
von: Schleifer, Abigail Victoria Gurin, et al.
Veröffentlicht: (2026)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
von: Ehsan, Md. Alvee, et al.
Veröffentlicht: (2025)
von: Ehsan, Md. Alvee, et al.
Veröffentlicht: (2025)
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning
von: Li, Yafu, et al.
Veröffentlicht: (2025)
von: Li, Yafu, et al.
Veröffentlicht: (2025)
Guiding and Diversifying LLM-Based Story Generation via Answer Set Programming
von: Wang, Phoebe J., et al.
Veröffentlicht: (2024)
von: Wang, Phoebe J., et al.
Veröffentlicht: (2024)
Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
von: Radliński, Łukasz, et al.
Veröffentlicht: (2025)
von: Radliński, Łukasz, et al.
Veröffentlicht: (2025)
QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform
von: Moses, Movina, et al.
Veröffentlicht: (2025)
von: Moses, Movina, et al.
Veröffentlicht: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
von: Kim, Wonjoong, et al.
Veröffentlicht: (2025)
von: Kim, Wonjoong, et al.
Veröffentlicht: (2025)
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
von: Kahana, Adar, et al.
Veröffentlicht: (2024)
von: Kahana, Adar, et al.
Veröffentlicht: (2024)
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation
von: Berlin, Konstantin, et al.
Veröffentlicht: (2026)
von: Berlin, Konstantin, et al.
Veröffentlicht: (2026)
Generative Language Models with Retrieval Augmented Generation for Automated Short Answer Scoring
von: Wang, Zifan, et al.
Veröffentlicht: (2024)
von: Wang, Zifan, et al.
Veröffentlicht: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
von: Shi, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Shi, Xiaofeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2025) -
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
von: Kiulian, Artur, et al.
Veröffentlicht: (2024) -
Dialectical Behavior Therapy Approach to LLM Prompting
von: Vitman, Oxana, et al.
Veröffentlicht: (2024) -
Towards LLM-based Autograding for Short Textual Answers
von: Schneider, Johannes, et al.
Veröffentlicht: (2023) -
Integrated Framework for LLM Evaluation with Answer Generation
von: Lee, Sujeong, et al.
Veröffentlicht: (2025)