Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Hasan, Mohammed Rakibul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
von: Zhang, Luyan, et al.
Veröffentlicht: (2025)
von: Zhang, Luyan, et al.
Veröffentlicht: (2025)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
von: Han, Lifeng, et al.
Veröffentlicht: (2016)
von: Han, Lifeng, et al.
Veröffentlicht: (2016)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
von: Hashemi, Helia, et al.
Veröffentlicht: (2024)
von: Hashemi, Helia, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
von: Hartnett, Peter, et al.
Veröffentlicht: (2026)
von: Hartnett, Peter, et al.
Veröffentlicht: (2026)
Improving the Capabilities of Large Language Model Based Marketing Analytics Copilots With Semantic Search And Fine-Tuning
von: Gao, Yilin, et al.
Veröffentlicht: (2024)
von: Gao, Yilin, et al.
Veröffentlicht: (2024)
BLT: Can Large Language Models Handle Basic Legal Text?
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2023)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2023)
Towards Conditioning Clinical Text Generation for User Control
von: Koraş, Osman Alperen, et al.
Veröffentlicht: (2025)
von: Koraş, Osman Alperen, et al.
Veröffentlicht: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
von: Cherif, Ahmed
Veröffentlicht: (2026)
von: Cherif, Ahmed
Veröffentlicht: (2026)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
ReTreVal: Reasoning Tree with Validation -- A Hybrid Framework for Enhanced LLM Multi-Step Reasoning
von: HS, Abhishek, et al.
Veröffentlicht: (2026)
von: HS, Abhishek, et al.
Veröffentlicht: (2026)
The Need for Guardrails with Large Language Models in Medical Safety-Critical Settings: An Artificial Intelligence Application in the Pharmacovigilance Ecosystem
von: Hakim, Joe B, et al.
Veröffentlicht: (2024)
von: Hakim, Joe B, et al.
Veröffentlicht: (2024)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
How to Evaluate Medical AI
von: Kopanichuk, Ilia, et al.
Veröffentlicht: (2025)
von: Kopanichuk, Ilia, et al.
Veröffentlicht: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
von: Rodriguez, David, et al.
Veröffentlicht: (2025)
von: Rodriguez, David, et al.
Veröffentlicht: (2025)
Scaling Laws for State Dynamics in Large Language Models
von: Li, Jacob X, et al.
Veröffentlicht: (2025)
von: Li, Jacob X, et al.
Veröffentlicht: (2025)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
von: Dayarathne, Ranul, et al.
Veröffentlicht: (2025)
von: Dayarathne, Ranul, et al.
Veröffentlicht: (2025)
Teaching a Language Model to Speak the Language of Tools
von: Emanuilov, Simeon
Veröffentlicht: (2025)
von: Emanuilov, Simeon
Veröffentlicht: (2025)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
von: Osborne, Philip, et al.
Veröffentlicht: (2025)
von: Osborne, Philip, et al.
Veröffentlicht: (2025)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
von: Karpurapu, Shanthi, et al.
Veröffentlicht: (2024)
von: Karpurapu, Shanthi, et al.
Veröffentlicht: (2024)
AskSport: Web Application for Sports Question-Answering
von: Onofre, Enzo B, et al.
Veröffentlicht: (2025)
von: Onofre, Enzo B, et al.
Veröffentlicht: (2025)
Challenges and Opportunities of NLP for HR Applications: A Discussion Paper
von: Leidner, Jochen L., et al.
Veröffentlicht: (2024)
von: Leidner, Jochen L., et al.
Veröffentlicht: (2024)
Comparative Analysis of AI Agent Architectures for Entity Relationship Classification
von: Berijanian, Maryam, et al.
Veröffentlicht: (2025)
von: Berijanian, Maryam, et al.
Veröffentlicht: (2025)
STLLM-DF: A Spatial-Temporal Large Language Model with Diffusion for Enhanced Multi-Mode Traffic System Forecasting
von: Shao, Zhiqi, et al.
Veröffentlicht: (2024)
von: Shao, Zhiqi, et al.
Veröffentlicht: (2024)
Open-TI: Open Traffic Intelligence with Augmented Language Model
von: Da, Longchao, et al.
Veröffentlicht: (2023)
von: Da, Longchao, et al.
Veröffentlicht: (2023)
AutoTRIZ: Automating Engineering Innovation with TRIZ and Large Language Models
von: Jiang, Shuo, et al.
Veröffentlicht: (2024)
von: Jiang, Shuo, et al.
Veröffentlicht: (2024)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
von: Ghandi, Taraneh, et al.
Veröffentlicht: (2026)
von: Ghandi, Taraneh, et al.
Veröffentlicht: (2026)
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
von: Sauter, Andreas, et al.
Veröffentlicht: (2026)
von: Sauter, Andreas, et al.
Veröffentlicht: (2026)
Mitigating Structural Noise in Low-Resource S2TT: An Optimized Cascaded Nepali-English Pipeline with Punctuation Restoration
von: Chongbang, Tangsang, et al.
Veröffentlicht: (2026)
von: Chongbang, Tangsang, et al.
Veröffentlicht: (2026)
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
von: Reddy, Sandeep, et al.
Veröffentlicht: (2025)
von: Reddy, Sandeep, et al.
Veröffentlicht: (2025)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
Large Language Models for Combinatorial Optimization of Design Structure Matrix
von: Jiang, Shuo, et al.
Veröffentlicht: (2024)
von: Jiang, Shuo, et al.
Veröffentlicht: (2024)
Exploring the Structure of AI-Induced Language Change in Scientific English
von: Galpin, Riley, et al.
Veröffentlicht: (2025)
von: Galpin, Riley, et al.
Veröffentlicht: (2025)
AI-Powered Detection of Inappropriate Language in Medical School Curricula
von: Salavati, Chiman, et al.
Veröffentlicht: (2025)
von: Salavati, Chiman, et al.
Veröffentlicht: (2025)
OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
von: Tao, Xinli, et al.
Veröffentlicht: (2025)
von: Tao, Xinli, et al.
Veröffentlicht: (2025)
VERITAS-NLI : Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference
von: Shah, Arjun, et al.
Veröffentlicht: (2024)
von: Shah, Arjun, et al.
Veröffentlicht: (2024)
A Llama walks into the 'Bar': Efficient Supervised Fine-Tuning for Legal Reasoning in the Multi-state Bar Exam
von: Fernandes, Rean, et al.
Veröffentlicht: (2025)
von: Fernandes, Rean, et al.
Veröffentlicht: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
Introducing Brain-like Concepts to Embodied Hand-crafted Dialog Management System
von: Joublin, Frank, et al.
Veröffentlicht: (2024)
von: Joublin, Frank, et al.
Veröffentlicht: (2024)
Project Riley: Multimodal Multi-Agent LLM Collaboration with Emotional Reasoning and Voting
von: Ortigoso, Ana Rita, et al.
Veröffentlicht: (2025)
von: Ortigoso, Ana Rita, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
von: Zhang, Luyan, et al.
Veröffentlicht: (2025) -
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
von: Han, Lifeng, et al.
Veröffentlicht: (2016) -
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
von: Hashemi, Helia, et al.
Veröffentlicht: (2024) -
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
von: Hartnett, Peter, et al.
Veröffentlicht: (2026) -
Improving the Capabilities of Large Language Model Based Marketing Analytics Copilots With Semantic Search And Fine-Tuning
von: Gao, Yilin, et al.
Veröffentlicht: (2024)