Benchmarking ChatGPT on Algorithmic Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | McLeish, Sean, Schwarzschild, Avi, Goldstein, Tom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The CLRS-Text Algorithmic Reasoning Language Benchmark
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
Transformers Can Do Arithmetic with the Right Embeddings
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
Fairness of ChatGPT
von: Li, Yunqi, et al.
Veröffentlicht: (2023)
von: Li, Yunqi, et al.
Veröffentlicht: (2023)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
von: Urchs, Stefanie, et al.
Veröffentlicht: (2023)
von: Urchs, Stefanie, et al.
Veröffentlicht: (2023)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
Can we trust the evaluation on ChatGPT?
von: Aiyappa, Rachith, et al.
Veröffentlicht: (2023)
von: Aiyappa, Rachith, et al.
Veröffentlicht: (2023)
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
von: Juzek, Tom S., et al.
Veröffentlicht: (2024)
von: Juzek, Tom S., et al.
Veröffentlicht: (2024)
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities
von: Chen, Yuhao, et al.
Veröffentlicht: (2023)
von: Chen, Yuhao, et al.
Veröffentlicht: (2023)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Evaluating the Performance of ChatGPT for Spam Email Detection
von: Si, Shijing, et al.
Veröffentlicht: (2024)
von: Si, Shijing, et al.
Veröffentlicht: (2024)
Jill Watson: A Virtual Teaching Assistant powered by ChatGPT
von: Taneja, Karan, et al.
Veröffentlicht: (2024)
von: Taneja, Karan, et al.
Veröffentlicht: (2024)
Is ChatGPT Transforming Academics' Writing Style?
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
Does ChatGPT Have a Mind?
von: Goldstein, Simon, et al.
Veröffentlicht: (2024)
von: Goldstein, Simon, et al.
Veröffentlicht: (2024)
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges
von: Huang, Yizheng, et al.
Veröffentlicht: (2024)
von: Huang, Yizheng, et al.
Veröffentlicht: (2024)
Transforming Dental Diagnostics with Artificial Intelligence: Advanced Integration of ChatGPT and Large Language Models for Patient Care
von: Nia, Masoumeh Farhadi, et al.
Veröffentlicht: (2024)
von: Nia, Masoumeh Farhadi, et al.
Veröffentlicht: (2024)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
von: Du, Shaoshuai, et al.
Veröffentlicht: (2025)
von: Du, Shaoshuai, et al.
Veröffentlicht: (2025)
ChatGPT Needs SPADE (Sustainability, PrivAcy, Digital divide, and Ethics) Evaluation: A Review
von: Khowaja, Sunder Ali, et al.
Veröffentlicht: (2023)
von: Khowaja, Sunder Ali, et al.
Veröffentlicht: (2023)
Leveraging Codebook Knowledge with NLI and ChatGPT for Zero-Shot Political Relation Classification
von: Hu, Yibo, et al.
Veröffentlicht: (2023)
von: Hu, Yibo, et al.
Veröffentlicht: (2023)
Large Language Models are Pattern Matchers: Editing Semi-Structured and Structured Documents with ChatGPT
von: Weber, Irene
Veröffentlicht: (2024)
von: Weber, Irene
Veröffentlicht: (2024)
ArXivBench: When You Should Avoid Using ChatGPT for Academic Writing
von: Li, Ning, et al.
Veröffentlicht: (2025)
von: Li, Ning, et al.
Veröffentlicht: (2025)
Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking
von: Bayani, David
Veröffentlicht: (2023)
von: Bayani, David
Veröffentlicht: (2023)
Detecting mental disorder on social media: a ChatGPT-augmented explainable approach
von: Belcastro, Loris, et al.
Veröffentlicht: (2024)
von: Belcastro, Loris, et al.
Veröffentlicht: (2024)
Can ChatGPT support software verification?
von: Janßen, Christian, et al.
Veröffentlicht: (2023)
von: Janßen, Christian, et al.
Veröffentlicht: (2023)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering Practice
von: Khojah, Ranim, et al.
Veröffentlicht: (2024)
von: Khojah, Ranim, et al.
Veröffentlicht: (2024)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
von: Iqbal, Umar, et al.
Veröffentlicht: (2023)
EHSAN: Leveraging ChatGPT in a Hybrid Framework for Arabic Aspect-Based Sentiment Analysis in Healthcare
von: Alamoudi, Eman, et al.
Veröffentlicht: (2025)
von: Alamoudi, Eman, et al.
Veröffentlicht: (2025)
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
von: Chen, Shan, et al.
Veröffentlicht: (2023)
von: Chen, Shan, et al.
Veröffentlicht: (2023)
Primacy Effect of ChatGPT
von: Wang, Yiwei, et al.
Veröffentlicht: (2023)
von: Wang, Yiwei, et al.
Veröffentlicht: (2023)
A "Perspectival" Mirror of the Elephant: Investigating Language Bias on Google, ChatGPT, YouTube, and Wikipedia
von: Luo, Queenie, et al.
Veröffentlicht: (2023)
von: Luo, Queenie, et al.
Veröffentlicht: (2023)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic
von: Al-Khalifa, Shahad, et al.
Veröffentlicht: (2024)
von: Al-Khalifa, Shahad, et al.
Veröffentlicht: (2024)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
von: Zhang, Juzheng, et al.
Veröffentlicht: (2025)
von: Zhang, Juzheng, et al.
Veröffentlicht: (2025)
ChatGPT for Conversational Recommendation: Refining Recommendations by Reprompting with Feedback
von: Spurlock, Kyle Dylan, et al.
Veröffentlicht: (2024)
von: Spurlock, Kyle Dylan, et al.
Veröffentlicht: (2024)
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
von: Mao, Rui, et al.
Veröffentlicht: (2023)
von: Mao, Rui, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The CLRS-Text Algorithmic Reasoning Language Benchmark
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024) -
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025) -
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
von: Lee, Sangyun, et al.
Veröffentlicht: (2026) -
Transformers Can Do Arithmetic with the Right Embeddings
von: McLeish, Sean, et al.
Veröffentlicht: (2024) -
Fairness of ChatGPT
von: Li, Yunqi, et al.
Veröffentlicht: (2023)