Salvato in:
| Autori principali: | , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.04069 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914963238420480 |
|---|---|
| author | Laskar, Md Tahmid Rahman Alqahtani, Sawsan Bari, M Saiful Rahman, Mizanur Khan, Mohammad Abdullah Matin Khan, Haidar Jahan, Israt Bhuiyan, Amran Tan, Chee Wei Parvez, Md Rizwan Hoque, Enamul Joty, Shafiq Huang, Jimmy |
| author_facet | Laskar, Md Tahmid Rahman Alqahtani, Sawsan Bari, M Saiful Rahman, Mizanur Khan, Mohammad Abdullah Matin Khan, Haidar Jahan, Israt Bhuiyan, Amran Tan, Chee Wei Parvez, Md Rizwan Hoque, Enamul Joty, Shafiq Huang, Jimmy |
| contents | Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them in real-world applications to ensure they produce reliable performance. Despite the well-established importance of evaluating LLMs in the community, the complexity of the evaluation process has led to varied evaluation setups, causing inconsistencies in findings and interpretations. To address this, we systematically review the primary challenges and limitations causing these inconsistencies and unreliable evaluations in various steps of LLM evaluation. Based on our critical review, we present our perspectives and recommendations to ensure LLM evaluations are reproducible, reliable, and robust. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_04069 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations Laskar, Md Tahmid Rahman Alqahtani, Sawsan Bari, M Saiful Rahman, Mizanur Khan, Mohammad Abdullah Matin Khan, Haidar Jahan, Israt Bhuiyan, Amran Tan, Chee Wei Parvez, Md Rizwan Hoque, Enamul Joty, Shafiq Huang, Jimmy Computation and Language Artificial Intelligence Machine Learning Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them in real-world applications to ensure they produce reliable performance. Despite the well-established importance of evaluating LLMs in the community, the complexity of the evaluation process has led to varied evaluation setups, causing inconsistencies in findings and interpretations. To address this, we systematically review the primary challenges and limitations causing these inconsistencies and unreliable evaluations in various steps of LLM evaluation. Based on our critical review, we present our perspectives and recommendations to ensure LLM evaluations are reproducible, reliable, and robust. |
| title | A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2407.04069 |