Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kostiuk, Yevhen, Vitman, Oxana, Gagała, Łukasz, Kiulian, Artur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2025)
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2025)
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
Dialectical Behavior Therapy Approach to LLM Prompting
von: Vitman, Oxana, et al.
Veröffentlicht: (2024)
von: Vitman, Oxana, et al.
Veröffentlicht: (2024)
Towards Multilingual LLM Evaluation for European Languages
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024)
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024)
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2026)
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2026)
Implementing a Nordic-Baltic Federated Health Data Network: a case report
von: Chomutare, Taridzo, et al.
Veröffentlicht: (2024)
von: Chomutare, Taridzo, et al.
Veröffentlicht: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2025)
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2025)
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers
von: Dhakal, Prakash, et al.
Veröffentlicht: (2024)
von: Dhakal, Prakash, et al.
Veröffentlicht: (2024)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2025)
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2025)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
von: Gupta, Ashray, et al.
Veröffentlicht: (2025)
von: Gupta, Ashray, et al.
Veröffentlicht: (2025)
Measuring Moral LLM Responses in Multilingual Capacities
von: Basu, Kimaya, et al.
Veröffentlicht: (2025)
von: Basu, Kimaya, et al.
Veröffentlicht: (2025)
M-Prometheus: A Suite of Open Multilingual LLM Judges
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Multilingual Large Language Models do not comprehend all natural languages to equal degrees
von: Moskvina, Natalia, et al.
Veröffentlicht: (2026)
von: Moskvina, Natalia, et al.
Veröffentlicht: (2026)
Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation
von: Mohamed, Asim, et al.
Veröffentlicht: (2025)
von: Mohamed, Asim, et al.
Veröffentlicht: (2025)
MELA: Multilingual Evaluation of Linguistic Acceptability
von: Zhang, Ziyin, et al.
Veröffentlicht: (2023)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2023)
Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
von: Radliński, Łukasz, et al.
Veröffentlicht: (2025)
von: Radliński, Łukasz, et al.
Veröffentlicht: (2025)
RDF-Based Structured Quality Assessment Representation of Multilingual LLM Evaluations
von: Gwozdz, Jonas, et al.
Veröffentlicht: (2025)
von: Gwozdz, Jonas, et al.
Veröffentlicht: (2025)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
von: Sitaram, Sunayana, et al.
Veröffentlicht: (2025)
von: Sitaram, Sunayana, et al.
Veröffentlicht: (2025)
Towards Safe Multilingual Frontier AI
von: Kanepajs, Artūrs, et al.
Veröffentlicht: (2024)
von: Kanepajs, Artūrs, et al.
Veröffentlicht: (2024)
M-IFEval: Multilingual Instruction-Following Evaluation
von: Dussolle, Antoine, et al.
Veröffentlicht: (2025)
von: Dussolle, Antoine, et al.
Veröffentlicht: (2025)
The Roles of English in Evaluating Multilingual Language Models
von: Poelman, Wessel, et al.
Veröffentlicht: (2024)
von: Poelman, Wessel, et al.
Veröffentlicht: (2024)
Translation as a Scalable Proxy for Multilingual Evaluation
von: Issaka, Sheriff, et al.
Veröffentlicht: (2026)
von: Issaka, Sheriff, et al.
Veröffentlicht: (2026)
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
von: Khan, Shaharukh, et al.
Veröffentlicht: (2026)
von: Khan, Shaharukh, et al.
Veröffentlicht: (2026)
FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
von: Munis, Evren Ayberk, et al.
Veröffentlicht: (2025)
von: Munis, Evren Ayberk, et al.
Veröffentlicht: (2025)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
von: Chen, Pinzhen, et al.
Veröffentlicht: (2024)
von: Chen, Pinzhen, et al.
Veröffentlicht: (2024)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
von: Chaudhary, Manav, et al.
Veröffentlicht: (2024)
von: Chaudhary, Manav, et al.
Veröffentlicht: (2024)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
von: Wu, Wanxing, et al.
Veröffentlicht: (2026)
von: Wu, Wanxing, et al.
Veröffentlicht: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages
von: Li, Haolin, et al.
Veröffentlicht: (2025)
von: Li, Haolin, et al.
Veröffentlicht: (2025)
Multilingual Training and Evaluation Resources for Vision-Language Models
von: Baiamonte, Daniela, et al.
Veröffentlicht: (2026)
von: Baiamonte, Daniela, et al.
Veröffentlicht: (2026)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
GPT-SW3: An Autoregressive Language Model for the Nordic Languages
von: Ekgren, Ariel, et al.
Veröffentlicht: (2023)
von: Ekgren, Ariel, et al.
Veröffentlicht: (2023)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
von: Min, Hyangsuk, et al.
Veröffentlicht: (2025)
von: Min, Hyangsuk, et al.
Veröffentlicht: (2025)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
von: Tao, Zhen, et al.
Veröffentlicht: (2024)
von: Tao, Zhen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
von: Kostiuk, Yevhen, et al.
Veröffentlicht: (2025) -
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
von: Kiulian, Artur, et al.
Veröffentlicht: (2024) -
Dialectical Behavior Therapy Approach to LLM Prompting
von: Vitman, Oxana, et al.
Veröffentlicht: (2024) -
Towards Multilingual LLM Evaluation for European Languages
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024) -
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
von: Kiulian, Artur, et al.
Veröffentlicht: (2024)