Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Alam, Firoj, Bhatia, Gagan, Laskar, Sahinur Rahman, Chowdhury, Shammur Absar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
di: Alam, Firoj, et al.
Pubblicazione: (2025)
di: Alam, Firoj, et al.
Pubblicazione: (2025)
NativQA Framework: Enabling LLMs and VLMs with Native, Local, and Everyday Knowledge
di: Alam, Firoj, et al.
Pubblicazione: (2025)
di: Alam, Firoj, et al.
Pubblicazione: (2025)
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs
di: Hasan, Md. Arid, et al.
Pubblicazione: (2024)
di: Hasan, Md. Arid, et al.
Pubblicazione: (2024)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2025)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2025)
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
di: Bhatti, Hunzalah Hassan, et al.
Pubblicazione: (2026)
di: Bhatti, Hunzalah Hassan, et al.
Pubblicazione: (2026)
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2024)
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2024)
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
di: Alam, Firoj, et al.
Pubblicazione: (2026)
di: Alam, Firoj, et al.
Pubblicazione: (2026)
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
di: Sukhadia, Vrunda N., et al.
Pubblicazione: (2026)
di: Sukhadia, Vrunda N., et al.
Pubblicazione: (2026)
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
di: Ersoy, Asım, et al.
Pubblicazione: (2025)
di: Ersoy, Asım, et al.
Pubblicazione: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025)
di: Collot, Stephane, et al.
Pubblicazione: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
di: Gupta, Raavi, et al.
Pubblicazione: (2025)
di: Gupta, Raavi, et al.
Pubblicazione: (2025)
WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
di: Ali, Zien Sheikh, et al.
Pubblicazione: (2026)
di: Ali, Zien Sheikh, et al.
Pubblicazione: (2026)
MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
di: Ali, Zien Sheikh, et al.
Pubblicazione: (2026)
di: Ali, Zien Sheikh, et al.
Pubblicazione: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Towards Multilingual LLM Evaluation for European Languages
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
di: García-Ferrero, Iker, et al.
Pubblicazione: (2024)
di: García-Ferrero, Iker, et al.
Pubblicazione: (2024)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
di: Ye, Ziyi, et al.
Pubblicazione: (2024)
di: Ye, Ziyi, et al.
Pubblicazione: (2024)
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset
di: Chowdhury, Mohammed Nowshad Ruhani, et al.
Pubblicazione: (2026)
di: Chowdhury, Mohammed Nowshad Ruhani, et al.
Pubblicazione: (2026)
Automated Rewards via LLM-Generated Progress Functions
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2024)
di: Sarukkai, Vishnu, et al.
Pubblicazione: (2024)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
di: Muhamed, Aashiq
Pubblicazione: (2025)
di: Muhamed, Aashiq
Pubblicazione: (2025)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
di: Haussmann, Aden
Pubblicazione: (2025)
di: Haussmann, Aden
Pubblicazione: (2025)
GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge
di: Chowdhury, Shammur Absar, et al.
Pubblicazione: (2024)
di: Chowdhury, Shammur Absar, et al.
Pubblicazione: (2024)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
di: Duo, Jiangshan, et al.
Pubblicazione: (2026)
di: Duo, Jiangshan, et al.
Pubblicazione: (2026)
Investigating Non-Transitivity in LLM-as-a-Judge
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
di: Geng, Mingmeng, et al.
Pubblicazione: (2025)
di: Geng, Mingmeng, et al.
Pubblicazione: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
di: Ali, Mehdi, et al.
Pubblicazione: (2025)
di: Ali, Mehdi, et al.
Pubblicazione: (2025)
TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation
di: Grimal, Paul, et al.
Pubblicazione: (2023)
di: Grimal, Paul, et al.
Pubblicazione: (2023)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models
di: Mousi, Basel, et al.
Pubblicazione: (2026)
di: Mousi, Basel, et al.
Pubblicazione: (2026)
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
di: Singh, Janvijay, et al.
Pubblicazione: (2025)
di: Singh, Janvijay, et al.
Pubblicazione: (2025)
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
di: Agarwal, Mehul, et al.
Pubblicazione: (2026)
di: Agarwal, Mehul, et al.
Pubblicazione: (2026)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026)
di: Yang, Jinming, et al.
Pubblicazione: (2026)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
di: Gera, Ariel, et al.
Pubblicazione: (2024)
di: Gera, Ariel, et al.
Pubblicazione: (2024)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
di: Salinas, David, et al.
Pubblicazione: (2025)
di: Salinas, David, et al.
Pubblicazione: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025)
di: Clegg, Kester, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
di: Alam, Firoj, et al.
Pubblicazione: (2025) -
NativQA Framework: Enabling LLMs and VLMs with Native, Local, and Everyday Knowledge
di: Alam, Firoj, et al.
Pubblicazione: (2025) -
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs
di: Hasan, Md. Arid, et al.
Pubblicazione: (2024) -
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2025) -
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
di: Bhatti, Hunzalah Hassan, et al.
Pubblicazione: (2026)