The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic
Fuente:
arXiv
Salvato in:
| Autori principali: | Al-Khalifa, Shahad, Al-Khalifa, Hend |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Survey of Large Language Models for Arabic Language and its Dialects
di: Mashaabi, Malak, et al.
Pubblicazione: (2024)
di: Mashaabi, Malak, et al.
Pubblicazione: (2024)
The Landscape of Arabic Large Language Models (ALLMs): A New Era for Arabic Language Technology
di: Al-Khalifa, Shahad, et al.
Pubblicazione: (2025)
di: Al-Khalifa, Shahad, et al.
Pubblicazione: (2025)
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
di: Panda, Srikant, et al.
Pubblicazione: (2025)
di: Panda, Srikant, et al.
Pubblicazione: (2025)
Gender Stereotypes in Professional Roles Among Saudis: An Analytical Study of AI-Generated Images Using Language Models
di: AlKhalifah, Khaloud S., et al.
Pubblicazione: (2025)
di: AlKhalifah, Khaloud S., et al.
Pubblicazione: (2025)
From Code-Centric to Concept-Centric: Teaching NLP with LLM-Assisted "Vibe Coding"
di: Al-Khalifa, Hend
Pubblicazione: (2026)
di: Al-Khalifa, Hend
Pubblicazione: (2026)
GLARE: Google Apps Arabic Reviews Dataset
di: AlGhamdi, Fatima, et al.
Pubblicazione: (2024)
di: AlGhamdi, Fatima, et al.
Pubblicazione: (2024)
Benchmarking ChatGPT on Algorithmic Reasoning
di: McLeish, Sean, et al.
Pubblicazione: (2024)
di: McLeish, Sean, et al.
Pubblicazione: (2024)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Can ChatGPT Really Understand Modern Chinese Poetry?
di: Wang, Shanshan, et al.
Pubblicazione: (2026)
di: Wang, Shanshan, et al.
Pubblicazione: (2026)
Primacy Effect of ChatGPT
di: Wang, Yiwei, et al.
Pubblicazione: (2023)
di: Wang, Yiwei, et al.
Pubblicazione: (2023)
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities
di: Chen, Yuhao, et al.
Pubblicazione: (2023)
di: Chen, Yuhao, et al.
Pubblicazione: (2023)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
di: Van Long, Phuoc Pham, et al.
Pubblicazione: (2023)
di: Van Long, Phuoc Pham, et al.
Pubblicazione: (2023)
CLEANANERCorp: Identifying and Correcting Incorrect Labels in the ANERcorp Dataset
di: Al-Duwais, Mashael, et al.
Pubblicazione: (2024)
di: Al-Duwais, Mashael, et al.
Pubblicazione: (2024)
Tibyan Corpus: Balanced and Comprehensive Error Coverage Corpus Using ChatGPT for Arabic Grammatical Error Correction
di: Alrehili, Ahlam, et al.
Pubblicazione: (2024)
di: Alrehili, Ahlam, et al.
Pubblicazione: (2024)
ChatGPT Alternative Solutions: Large Language Models Survey
di: Alipour, Hanieh, et al.
Pubblicazione: (2024)
di: Alipour, Hanieh, et al.
Pubblicazione: (2024)
Fairness of ChatGPT
di: Li, Yunqi, et al.
Pubblicazione: (2023)
di: Li, Yunqi, et al.
Pubblicazione: (2023)
Emojis Decoded: Leveraging ChatGPT for Enhanced Understanding in Social Media Communications
di: Zhou, Yuhang, et al.
Pubblicazione: (2024)
di: Zhou, Yuhang, et al.
Pubblicazione: (2024)
ADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational Sociopragmatics
di: Al-Khalifa, Hend, et al.
Pubblicazione: (2026)
di: Al-Khalifa, Hend, et al.
Pubblicazione: (2026)
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
di: Mao, Rui, et al.
Pubblicazione: (2023)
di: Mao, Rui, et al.
Pubblicazione: (2023)
ChatGPT as speechwriter for the French presidents
di: Labbé, Dominique, et al.
Pubblicazione: (2024)
di: Labbé, Dominique, et al.
Pubblicazione: (2024)
The Human and the Mechanical: logos, truthfulness, and ChatGPT
di: Giannakidou, Anastasia, et al.
Pubblicazione: (2024)
di: Giannakidou, Anastasia, et al.
Pubblicazione: (2024)
Does ChatGPT Have a Mind?
di: Goldstein, Simon, et al.
Pubblicazione: (2024)
di: Goldstein, Simon, et al.
Pubblicazione: (2024)
On Prompt Sensitivity of ChatGPT in Affective Computing
di: Amin, Mostafa M., et al.
Pubblicazione: (2024)
di: Amin, Mostafa M., et al.
Pubblicazione: (2024)
A Survey on the Real Power of ChatGPT
di: Liu, Ming, et al.
Pubblicazione: (2024)
di: Liu, Ming, et al.
Pubblicazione: (2024)
Can ChatGPT Learn to Count Letters?
di: Conde, Javier, et al.
Pubblicazione: (2025)
di: Conde, Javier, et al.
Pubblicazione: (2025)
ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
di: Tu, Shangqing, et al.
Pubblicazione: (2023)
di: Tu, Shangqing, et al.
Pubblicazione: (2023)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
di: Hijazi, Faris, et al.
Pubblicazione: (2024)
di: Hijazi, Faris, et al.
Pubblicazione: (2024)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
di: Urchs, Stefanie, et al.
Pubblicazione: (2023)
di: Urchs, Stefanie, et al.
Pubblicazione: (2023)
MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection
di: Al-Henaki, Lubna, et al.
Pubblicazione: (2025)
di: Al-Henaki, Lubna, et al.
Pubblicazione: (2025)
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
di: Li, Lingyao, et al.
Pubblicazione: (2023)
di: Li, Lingyao, et al.
Pubblicazione: (2023)
What is the Best Way for ChatGPT to Translate Poetry?
di: Wang, Shanshan, et al.
Pubblicazione: (2024)
di: Wang, Shanshan, et al.
Pubblicazione: (2024)
Evaluating ChatGPT on Nuclear Domain-Specific Data
di: Anwar, Muhammad, et al.
Pubblicazione: (2024)
di: Anwar, Muhammad, et al.
Pubblicazione: (2024)
Demystifying ChatGPT: How It Masters Genre Recognition
di: Raj, Subham, et al.
Pubblicazione: (2025)
di: Raj, Subham, et al.
Pubblicazione: (2025)
Unmasking the giant: A comprehensive evaluation of ChatGPT's proficiency in coding algorithms and data structures
di: Arefin, Sayed Erfan, et al.
Pubblicazione: (2023)
di: Arefin, Sayed Erfan, et al.
Pubblicazione: (2023)
BALSAM: A Platform for Benchmarking Arabic Large Language Models
di: Al-Matham, Rawan, et al.
Pubblicazione: (2025)
di: Al-Matham, Rawan, et al.
Pubblicazione: (2025)
LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding
di: Jubair, Sheikh, et al.
Pubblicazione: (2025)
di: Jubair, Sheikh, et al.
Pubblicazione: (2025)
ALARB: An Arabic Legal Argument Reasoning Benchmark
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
AuditGPT: Auditing Smart Contracts with ChatGPT
di: Xia, Shihao, et al.
Pubblicazione: (2024)
di: Xia, Shihao, et al.
Pubblicazione: (2024)
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
di: Huang, Fan, et al.
Pubblicazione: (2024)
di: Huang, Fan, et al.
Pubblicazione: (2024)
Exploiting ChatGPT for Diagnosing Autism-Associated Language Disorders and Identifying Distinct Features
di: Hu, Chuanbo, et al.
Pubblicazione: (2024)
di: Hu, Chuanbo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Survey of Large Language Models for Arabic Language and its Dialects
di: Mashaabi, Malak, et al.
Pubblicazione: (2024) -
The Landscape of Arabic Large Language Models (ALLMs): A New Era for Arabic Language Technology
di: Al-Khalifa, Shahad, et al.
Pubblicazione: (2025) -
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
di: Panda, Srikant, et al.
Pubblicazione: (2025) -
Gender Stereotypes in Professional Roles Among Saudis: An Analytical Study of AI-Generated Images Using Language Models
di: AlKhalifah, Khaloud S., et al.
Pubblicazione: (2025) -
From Code-Centric to Concept-Centric: Teaching NLP with LLM-Assisted "Vibe Coding"
di: Al-Khalifa, Hend
Pubblicazione: (2026)