Can we trust the evaluation on ChatGPT?
Fuente:
arXiv
Saved in:
| Main Authors: | Aiyappa, Rachith, An, Jisun, Kwak, Haewoon, Ahn, Yong-Yeol |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking zero-shot stance detection with FlanT5-XXL: Insights from training data, prompting, and decoding strategies into its near-SoTA performance
by: Aiyappa, Rachith, et al.
Published: (2024)
by: Aiyappa, Rachith, et al.
Published: (2024)
A semantic embedding space based on large language models for modelling human beliefs
by: Lee, Byunghwee, et al.
Published: (2024)
by: Lee, Byunghwee, et al.
Published: (2024)
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
by: Huang, Fan, et al.
Published: (2024)
by: Huang, Fan, et al.
Published: (2024)
What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
by: Malone, Joseph, et al.
Published: (2025)
by: Malone, Joseph, et al.
Published: (2025)
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
by: Muralidharan, Rasika, et al.
Published: (2025)
by: Muralidharan, Rasika, et al.
Published: (2025)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
ToBlend: Token-Level Blending With an Ensemble of LLMs to Attack AI-Generated Text Detection
by: Huang, Fan, et al.
Published: (2024)
by: Huang, Fan, et al.
Published: (2024)
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023)
by: Li, Yunqi, et al.
Published: (2023)
Benchmarking ChatGPT on Algorithmic Reasoning
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities
by: Chen, Yuhao, et al.
Published: (2023)
by: Chen, Yuhao, et al.
Published: (2023)
Evaluating the Performance of ChatGPT for Spam Email Detection
by: Si, Shijing, et al.
Published: (2024)
by: Si, Shijing, et al.
Published: (2024)
Jill Watson: A Virtual Teaching Assistant powered by ChatGPT
by: Taneja, Karan, et al.
Published: (2024)
by: Taneja, Karan, et al.
Published: (2024)
Is ChatGPT Transforming Academics' Writing Style?
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
Can ChatGPT support software verification?
by: Janßen, Christian, et al.
Published: (2023)
by: Janßen, Christian, et al.
Published: (2023)
LLMs Can Infer Political Alignment from Online Conversations
by: Lee, Byunghwee, et al.
Published: (2026)
by: Lee, Byunghwee, et al.
Published: (2026)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges
by: Huang, Yizheng, et al.
Published: (2024)
by: Huang, Yizheng, et al.
Published: (2024)
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
by: Juzek, Tom S., et al.
Published: (2024)
by: Juzek, Tom S., et al.
Published: (2024)
Transforming Dental Diagnostics with Artificial Intelligence: Advanced Integration of ChatGPT and Large Language Models for Patient Care
by: Nia, Masoumeh Farhadi, et al.
Published: (2024)
by: Nia, Masoumeh Farhadi, et al.
Published: (2024)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025)
by: Du, Shaoshuai, et al.
Published: (2025)
Can ChatGPT Learn to Count Letters?
by: Conde, Javier, et al.
Published: (2025)
by: Conde, Javier, et al.
Published: (2025)
Who can we trust? LLM-as-a-jury for Comparative Assessment
by: Qian, Mengjie, et al.
Published: (2026)
by: Qian, Mengjie, et al.
Published: (2026)
ChatGPT Needs SPADE (Sustainability, PrivAcy, Digital divide, and Ethics) Evaluation: A Review
by: Khowaja, Sunder Ali, et al.
Published: (2023)
by: Khowaja, Sunder Ali, et al.
Published: (2023)
Leveraging Codebook Knowledge with NLI and ChatGPT for Zero-Shot Political Relation Classification
by: Hu, Yibo, et al.
Published: (2023)
by: Hu, Yibo, et al.
Published: (2023)
Can ChatGPT Diagnose Alzheimer's Disease?
by: Nguyen, Quoc-Toan, et al.
Published: (2025)
by: Nguyen, Quoc-Toan, et al.
Published: (2025)
ArXivBench: When You Should Avoid Using ChatGPT for Academic Writing
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Large Language Models are Pattern Matchers: Editing Semi-Structured and Structured Documents with ChatGPT
by: Weber, Irene
Published: (2024)
by: Weber, Irene
Published: (2024)
Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking
by: Bayani, David
Published: (2023)
by: Bayani, David
Published: (2023)
Detecting mental disorder on social media: a ChatGPT-augmented explainable approach
by: Belcastro, Loris, et al.
Published: (2024)
by: Belcastro, Loris, et al.
Published: (2024)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
by: Kachwala, Zoher, et al.
Published: (2026)
by: Kachwala, Zoher, et al.
Published: (2026)
Can ChatGPT Really Understand Modern Chinese Poetry?
by: Wang, Shanshan, et al.
Published: (2026)
by: Wang, Shanshan, et al.
Published: (2026)
Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering Practice
by: Khojah, Ranim, et al.
Published: (2024)
by: Khojah, Ranim, et al.
Published: (2024)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
by: Iqbal, Umar, et al.
Published: (2023)
by: Iqbal, Umar, et al.
Published: (2023)
EHSAN: Leveraging ChatGPT in a Hybrid Framework for Arabic Aspect-Based Sentiment Analysis in Healthcare
by: Alamoudi, Eman, et al.
Published: (2025)
by: Alamoudi, Eman, et al.
Published: (2025)
Primacy Effect of ChatGPT
by: Wang, Yiwei, et al.
Published: (2023)
by: Wang, Yiwei, et al.
Published: (2023)
A "Perspectival" Mirror of the Elephant: Investigating Language Bias on Google, ChatGPT, YouTube, and Wikipedia
by: Luo, Queenie, et al.
Published: (2023)
by: Luo, Queenie, et al.
Published: (2023)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Similar Items
-
Benchmarking zero-shot stance detection with FlanT5-XXL: Insights from training data, prompting, and decoding strategies into its near-SoTA performance
by: Aiyappa, Rachith, et al.
Published: (2024) -
A semantic embedding space based on large language models for modelling human beliefs
by: Lee, Byunghwee, et al.
Published: (2024) -
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
by: Huang, Fan, et al.
Published: (2024) -
What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
by: Malone, Joseph, et al.
Published: (2025) -
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
by: Muralidharan, Rasika, et al.
Published: (2025)