Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pombal, José, Guerreiro, Nuno M., Rei, Ricardo, Martins, André F. T. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
von: Pombal, José, et al.
Veröffentlicht: (2026)
von: Pombal, José, et al.
Veröffentlicht: (2026)
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
von: Rei, Ricardo, et al.
Veröffentlicht: (2025)
von: Rei, Ricardo, et al.
Veröffentlicht: (2025)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
von: Farinhas, António, et al.
Veröffentlicht: (2025)
von: Farinhas, António, et al.
Veröffentlicht: (2025)
M-Prometheus: A Suite of Open Multilingual LLM Judges
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models on the Frame and Symbol Grounding Problems: A Zero-shot Benchmark
von: Oka, Shoko
Veröffentlicht: (2025)
von: Oka, Shoko
Veröffentlicht: (2025)
Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution
von: Jin, Qiao, et al.
Veröffentlicht: (2026)
von: Jin, Qiao, et al.
Veröffentlicht: (2026)
EuroLLM-9B: Technical Report
von: Martins, Pedro Henrique, et al.
Veröffentlicht: (2025)
von: Martins, Pedro Henrique, et al.
Veröffentlicht: (2025)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
EuroLLM-22B: Technical Report
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2026)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2026)
ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Models
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
A Scalable Framework for Evaluating Health Language Models
von: Mallinar, Neil, et al.
Veröffentlicht: (2025)
von: Mallinar, Neil, et al.
Veröffentlicht: (2025)
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
Using Zero-shot Prompting in the Automatic Creation and Expansion of Topic Taxonomies for Tagging Retail Banking Transactions
von: Moraes, Daniel de S., et al.
Veröffentlicht: (2024)
von: Moraes, Daniel de S., et al.
Veröffentlicht: (2024)
EuroBERT: Scaling Multilingual Encoders for European Languages
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
Meta-Reasoning Improves Tool Use in Large Language Models
von: Alazraki, Lisa, et al.
Veröffentlicht: (2024)
von: Alazraki, Lisa, et al.
Veröffentlicht: (2024)
MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
von: Li, Zekun, et al.
Veröffentlicht: (2024)
von: Li, Zekun, et al.
Veröffentlicht: (2024)
ZeFaV: Boosting Large Language Models for Zero-shot Fact Verification
von: Luu, Son T., et al.
Veröffentlicht: (2024)
von: Luu, Son T., et al.
Veröffentlicht: (2024)
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
Zero-shot Graph Reasoning via Retrieval Augmented Framework with LLMs
von: Li, Hanqing, et al.
Veröffentlicht: (2025)
von: Li, Hanqing, et al.
Veröffentlicht: (2025)
xTower: A Multilingual LLM for Explaining and Correcting Translation Errors
von: Treviso, Marcos, et al.
Veröffentlicht: (2024)
von: Treviso, Marcos, et al.
Veröffentlicht: (2024)
Zero-shot Cross-Lingual Transfer for Synthetic Data Generation in Grammatical Error Detection
von: Latouche, Gaetan Lopez, et al.
Veröffentlicht: (2024)
von: Latouche, Gaetan Lopez, et al.
Veröffentlicht: (2024)
CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models
von: Chun, Jon, et al.
Veröffentlicht: (2026)
von: Chun, Jon, et al.
Veröffentlicht: (2026)
LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
TALEC: Teach Your LLM to Evaluate in Specific Domain with In-house Criteria by Criteria Division and Zero-shot Plus Few-shot
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2024)
A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
von: Murthy, Rithesh, et al.
Veröffentlicht: (2025)
von: Murthy, Rithesh, et al.
Veröffentlicht: (2025)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
von: Alves, Duarte M., et al.
Veröffentlicht: (2024)
von: Alves, Duarte M., et al.
Veröffentlicht: (2024)
Effectiveness of Zero-shot-CoT in Japanese Prompts
von: Takayama, Shusuke, et al.
Veröffentlicht: (2025)
von: Takayama, Shusuke, et al.
Veröffentlicht: (2025)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models to Extract Information on Substance Use Disorder Severity from Clinical Notes: A Zero-shot Learning Approach
von: Mahbub, Maria, et al.
Veröffentlicht: (2024)
von: Mahbub, Maria, et al.
Veröffentlicht: (2024)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation
von: Haffoudhi, Samy, et al.
Veröffentlicht: (2026)
von: Haffoudhi, Samy, et al.
Veröffentlicht: (2026)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain
von: Obaidah, Qusai Abo, et al.
Veröffentlicht: (2024)
von: Obaidah, Qusai Abo, et al.
Veröffentlicht: (2024)
MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support
von: Farinhas, António, et al.
Veröffentlicht: (2026)
von: Farinhas, António, et al.
Veröffentlicht: (2026)
Generating Benchmarks for Factuality Evaluation of Language Models
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
von: Pombal, José, et al.
Veröffentlicht: (2025) -
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
von: Pombal, José, et al.
Veröffentlicht: (2026) -
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
von: Pombal, José, et al.
Veröffentlicht: (2025) -
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
von: Rei, Ricardo, et al.
Veröffentlicht: (2025) -
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
von: Farinhas, António, et al.
Veröffentlicht: (2025)