Improving Methodologies for LLM Evaluations Across Global Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Vij, Akriti, Chua, Benjamin, Ramiah, Darshini, Ng, En Qi, Morsidi, Mahran, Gangarapu, Naga Nikshith, Johnson, Sharmini, Wilfred, Vanessa, Kumaran, Vikneswaran, Lee, Wan Sie, Yang, Wenzhuo, Zheng, Yongsen, Black, Bill, Xia, Boming, Sun, Frank, Zhang, Hao, Lu, Qinghua, Ma, Suyu, Liu, Yue, Lo, Chi-kiu, Azadi, Fatemeh, Nejadgholi, Isar, Vajjala, Sowmya, Delaborde, Agnes, Rolin, Nicolas, Seimandi, Tom, Murakami, Akiko, Ishi, Haruto, Sekine, Satoshi, Semitsu, Takayuki, Sasaki, Tasuku, Kinuthia, Angela, Wangari, Jean, Michie, Michael, Kasaon, Stephanie, Baek, Hankyul, Noh, Jaewon, Nam, Kihyuk, Seo, Sang, Shin, Sungpil, Lee, Taewhi, Kim, Yongsu, Newbold-Harrop, Daisy, Wang, Jessica, Ghanem, Mahmoud, Hong, Vy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
by: Seah, Ee Wei, et al.
Published: (2026)
by: Seah, Ee Wei, et al.
Published: (2026)
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Gender-Neutral Machine Translation Strategies in Practice
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
The public employment service in the Republic of Korea
by: Sungpil Yang
Published: (2015)
by: Sungpil Yang
Published: (2015)
The Masked Effect™ Five Systemic Profiles of False Synergy in Organizational Performance
by: Ramiah, Pragalathan
Published: (2026)
by: Ramiah, Pragalathan
Published: (2026)
Decent work in Ahmedabad: an integrated approach
by: Darshini Mahadevia
Published: (2012)
by: Darshini Mahadevia
Published: (2012)
Pluralising Scholarship: Repositioning Doctor of Nursing Practice Faculty Through Boyer's Framework: A Discursive Paper
by: Rachel Wangari Kimani
Published: (2026)
by: Rachel Wangari Kimani
Published: (2026)
Raw data
by: Suman, Vajjala
Published: (2025)
by: Suman, Vajjala
Published: (2025)
IndicGEC: Powerful Models, or a Measurement Mirage?
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
The Problem with Safety Classification is not just the Models
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
Securing AI Systems: A Guide to Known Attacks and Impacts
by: Kiribuchi, Naoto, et al.
Published: (2025)
by: Kiribuchi, Naoto, et al.
Published: (2025)
Automated Analysis of Global AI Safety Initiatives: A Taxonomy-Driven LLM Approach
by: Semitsu, Takayuki, et al.
Published: (2026)
by: Semitsu, Takayuki, et al.
Published: (2026)
Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models
by: Lee, Kihyuk
Published: (2026)
by: Lee, Kihyuk
Published: (2026)
Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model
by: Lee, Kihyuk
Published: (2026)
by: Lee, Kihyuk
Published: (2026)
Unified Theory of Quartz Tuning Fork Resonators
by: Koh, Hankyul, et al.
Published: (2026)
by: Koh, Hankyul, et al.
Published: (2026)
Spinoza on Teleology, Action, and Explanatory Overdetermination
by: Stephen Harrop
Published: (2025)
by: Stephen Harrop
Published: (2025)
Sufficient Reason Vindicated
by: Stephen Harrop
Published: (2025)
by: Stephen Harrop
Published: (2025)
Confidence Intervals for the F1 Score: A Comparison of Four Methods
by: Lam, Kevin Fu Yuan, et al.
Published: (2023)
by: Lam, Kevin Fu Yuan, et al.
Published: (2023)
Understanding insect predator–prey interactions using camera trapping: A review of current research and perspectives
by: Gaëtan Seimandi‐Corda, et al.
Published: (2024)
by: Gaëtan Seimandi‐Corda, et al.
Published: (2024)
Superradiant Anti‐Stokes Fluorescence of Organic Dye J‐Aggregates
by: Hankyul Lee, et al.
Published: (2025)
by: Hankyul Lee, et al.
Published: (2025)
Dravidian language family through Universal Dependencies lens
by: Rama, Taraka, et al.
Published: (2024)
by: Rama, Taraka, et al.
Published: (2024)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Text Classification in the LLM Era -- Where do we stand?
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
by: Kamath, Gaurav, et al.
Published: (2025)
by: Kamath, Gaurav, et al.
Published: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
by: Simbeck, Katharina, et al.
Published: (2025)
by: Simbeck, Katharina, et al.
Published: (2025)
Determinants of Male Partner Participation in Antenatal Care at Kangundo Hospital, Kenya
by: Wangari Mutuku, et al.
Published: (2025)
by: Wangari Mutuku, et al.
Published: (2025)
Fast Quantum Convolutional Neural Networks for Low-Complexity Object Detection in Autonomous Driving Applications
by: Baek, Hankyul, et al.
Published: (2023)
by: Baek, Hankyul, et al.
Published: (2023)
A Dynamic Change of Microglial States Occurs During the Transition From Photoreceptor Degeneration to Regeneration in Zebrafish pde6c Mutants
by: Darshini Ravishankar, et al.
Published: (2026)
by: Darshini Ravishankar, et al.
Published: (2026)
Annotation Errors and NER: A Study with OntoNotes 5.0
by: Bernier-Colborne, Gabriel, et al.
Published: (2024)
by: Bernier-Colborne, Gabriel, et al.
Published: (2024)
Exports, skills, and wage inequality in Kenya's manufacturing firms
by: Bethuel Kinyanjui Kinuthia, et al.
Published: (2024)
by: Bethuel Kinyanjui Kinuthia, et al.
Published: (2024)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
LayerAct: Advanced Activation Mechanism for Robust Inference of CNNs
by: Yoon, Kihyuk, et al.
Published: (2023)
by: Yoon, Kihyuk, et al.
Published: (2023)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Break Out of a Pigeonhole: A Unified Framework for Examining Miscalibration, Bias, and Stereotype in Recommender Systems
by: Ahn, Yongsu, et al.
Published: (2023)
by: Ahn, Yongsu, et al.
Published: (2023)
Exploring Teachers' Perception of Artificial Intelligence: The Socio-emotional Deficiency as Opportunities and Challenges in Human-AI Complementarity in K-12 Education
by: Oh, Soon-young, et al.
Published: (2024)
by: Oh, Soon-young, et al.
Published: (2024)
Understanding Why ChatGPT Outperforms Humans in Visualization Design Advice
by: Ahn, Yongsu, et al.
Published: (2025)
by: Ahn, Yongsu, et al.
Published: (2025)
Similar Items
-
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
by: Seah, Ee Wei, et al.
Published: (2026) -
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
by: Dawkins, Hillary, et al.
Published: (2024) -
Gender-Neutral Machine Translation Strategies in Practice
by: Dawkins, Hillary, et al.
Published: (2025) -
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026) -
The public employment service in the Republic of Korea
by: Sungpil Yang
Published: (2015)