The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Kostiuk, Yevhen, Vitman, Oxana, Gagała, Łukasz, Kiulian, Artur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History
by: Kostiuk, Yevhen, et al.
Published: (2025)
by: Kostiuk, Yevhen, et al.
Published: (2025)
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
by: Kiulian, Artur, et al.
Published: (2024)
by: Kiulian, Artur, et al.
Published: (2024)
Dialectical Behavior Therapy Approach to LLM Prompting
by: Vitman, Oxana, et al.
Published: (2024)
by: Vitman, Oxana, et al.
Published: (2024)
Towards LLM-based Autograding for Short Textual Answers
by: Schneider, Johannes, et al.
Published: (2023)
by: Schneider, Johannes, et al.
Published: (2023)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
by: Kiulian, Artur, et al.
Published: (2024)
by: Kiulian, Artur, et al.
Published: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
by: Lai, Peichao, et al.
Published: (2025)
by: Lai, Peichao, et al.
Published: (2025)
Directed Graph-alignment Approach for Identification of Gaps in Short Answers
by: Sahu, Archana, et al.
Published: (2025)
by: Sahu, Archana, et al.
Published: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
by: Kostiuk, Yevhen, et al.
Published: (2026)
by: Kostiuk, Yevhen, et al.
Published: (2026)
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2024)
by: Li, Dawei, et al.
Published: (2024)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
by: Chu, Yucheng, et al.
Published: (2026)
by: Chu, Yucheng, et al.
Published: (2026)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
by: Chen, Luyu, et al.
Published: (2025)
by: Chen, Luyu, et al.
Published: (2025)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
by: Oh, Jungsuk, et al.
Published: (2025)
by: Oh, Jungsuk, et al.
Published: (2025)
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
by: Ye, Bingyang, et al.
Published: (2026)
by: Ye, Bingyang, et al.
Published: (2026)
Evaluating and Optimizing Educational Content with Large Language Model Judgments
by: He-Yueya, Joy, et al.
Published: (2024)
by: He-Yueya, Joy, et al.
Published: (2024)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
by: Henkel, Owen, et al.
Published: (2024)
by: Henkel, Owen, et al.
Published: (2024)
Beyond Scores: A Modular RAG-Based System for Automatic Short Answer Scoring with Feedback
by: Fateen, Menna, et al.
Published: (2024)
by: Fateen, Menna, et al.
Published: (2024)
ASAG2024: A Combined Benchmark for Short Answer Grading
by: Meyer, Gérôme, et al.
Published: (2024)
by: Meyer, Gérôme, et al.
Published: (2024)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
by: Cheng, Qi, et al.
Published: (2024)
by: Cheng, Qi, et al.
Published: (2024)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
by: Henkel, Owen, et al.
Published: (2023)
by: Henkel, Owen, et al.
Published: (2023)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
by: Aggarwal, Dishank, et al.
Published: (2024)
by: Aggarwal, Dishank, et al.
Published: (2024)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
by: Ji, Jiabao, et al.
Published: (2025)
by: Ji, Jiabao, et al.
Published: (2025)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
by: Polo, Felipe Maia, et al.
Published: (2025)
by: Polo, Felipe Maia, et al.
Published: (2025)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
by: Schleifer, Abigail Victoria Gurin, et al.
Published: (2026)
by: Schleifer, Abigail Victoria Gurin, et al.
Published: (2026)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
by: Ehsan, Md. Alvee, et al.
Published: (2025)
by: Ehsan, Md. Alvee, et al.
Published: (2025)
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning
by: Li, Yafu, et al.
Published: (2025)
by: Li, Yafu, et al.
Published: (2025)
Guiding and Diversifying LLM-Based Story Generation via Answer Set Programming
by: Wang, Phoebe J., et al.
Published: (2024)
by: Wang, Phoebe J., et al.
Published: (2024)
Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
by: Radliński, Łukasz, et al.
Published: (2025)
by: Radliński, Łukasz, et al.
Published: (2025)
QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform
by: Moses, Movina, et al.
Published: (2025)
by: Moses, Movina, et al.
Published: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
by: Zhou, Chengliang, et al.
Published: (2025)
by: Zhou, Chengliang, et al.
Published: (2025)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025)
by: Kim, Wonjoong, et al.
Published: (2025)
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
by: Kahana, Adar, et al.
Published: (2024)
by: Kahana, Adar, et al.
Published: (2024)
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation
by: Berlin, Konstantin, et al.
Published: (2026)
by: Berlin, Konstantin, et al.
Published: (2026)
Generative Language Models with Retrieval Augmented Generation for Automated Short Answer Scoring
by: Wang, Zifan, et al.
Published: (2024)
by: Wang, Zifan, et al.
Published: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
by: Shi, Xiaofeng, et al.
Published: (2025)
by: Shi, Xiaofeng, et al.
Published: (2025)
Similar Items
-
Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History
by: Kostiuk, Yevhen, et al.
Published: (2025) -
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
by: Kiulian, Artur, et al.
Published: (2024) -
Dialectical Behavior Therapy Approach to LLM Prompting
by: Vitman, Oxana, et al.
Published: (2024) -
Towards LLM-based Autograding for Short Textual Answers
by: Schneider, Johannes, et al.
Published: (2023) -
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)