ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
Fuente:
arXiv
Saved in:
| Main Authors: | Tu, Shangqing, Li, Chunyang, Yu, Jifan, Wang, Xiaozhi, Hou, Lei, Li, Juanzi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2023)
by: Tu, Shangqing, et al.
Published: (2023)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023)
by: Li, Yunqi, et al.
Published: (2023)
Primacy Effect of ChatGPT
by: Wang, Yiwei, et al.
Published: (2023)
by: Wang, Yiwei, et al.
Published: (2023)
DeepPrune: Parallel Scaling without Inter-trace Redundancy
by: Tu, Shangqing, et al.
Published: (2025)
by: Tu, Shangqing, et al.
Published: (2025)
A Survey on the Real Power of ChatGPT
by: Liu, Ming, et al.
Published: (2024)
by: Liu, Ming, et al.
Published: (2024)
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
by: Chen, Shan, et al.
Published: (2023)
by: Chen, Shan, et al.
Published: (2023)
Evaluating ChatGPT on Nuclear Domain-Specific Data
by: Anwar, Muhammad, et al.
Published: (2024)
by: Anwar, Muhammad, et al.
Published: (2024)
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
by: Li, Lingyao, et al.
Published: (2023)
by: Li, Lingyao, et al.
Published: (2023)
Automated Coding of Communication Data Using ChatGPT: Consistency Across Subgroups
by: Hao, Jiangang, et al.
Published: (2025)
by: Hao, Jiangang, et al.
Published: (2025)
AI and the Law: Evaluating ChatGPT's Performance in Legal Classification
by: Weichbroth, Pawel
Published: (2025)
by: Weichbroth, Pawel
Published: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
The Human and the Mechanical: logos, truthfulness, and ChatGPT
by: Giannakidou, Anastasia, et al.
Published: (2024)
by: Giannakidou, Anastasia, et al.
Published: (2024)
Does ChatGPT Have a Mind?
by: Goldstein, Simon, et al.
Published: (2024)
by: Goldstein, Simon, et al.
Published: (2024)
Can ChatGPT Learn to Count Letters?
by: Conde, Javier, et al.
Published: (2025)
by: Conde, Javier, et al.
Published: (2025)
On Prompt Sensitivity of ChatGPT in Affective Computing
by: Amin, Mostafa M., et al.
Published: (2024)
by: Amin, Mostafa M., et al.
Published: (2024)
What is the Best Way for ChatGPT to Translate Poetry?
by: Wang, Shanshan, et al.
Published: (2024)
by: Wang, Shanshan, et al.
Published: (2024)
ChatGPT as speechwriter for the French presidents
by: Labbé, Dominique, et al.
Published: (2024)
by: Labbé, Dominique, et al.
Published: (2024)
Benchmarking ChatGPT on Algorithmic Reasoning
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
by: Mao, Rui, et al.
Published: (2023)
by: Mao, Rui, et al.
Published: (2023)
Exploring ChatGPT and its Impact on Society
by: Haque, Md. Asraful, et al.
Published: (2024)
by: Haque, Md. Asraful, et al.
Published: (2024)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
Exploiting ChatGPT for Diagnosing Autism-Associated Language Disorders and Identifying Distinct Features
by: Hu, Chuanbo, et al.
Published: (2024)
by: Hu, Chuanbo, et al.
Published: (2024)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
by: Li, Victoria R., et al.
Published: (2024)
by: Li, Victoria R., et al.
Published: (2024)
Differentiate ChatGPT-generated and Human-written Medical Texts
by: Liao, Wenxiong, et al.
Published: (2023)
by: Liao, Wenxiong, et al.
Published: (2023)
AuditGPT: Auditing Smart Contracts with ChatGPT
by: Xia, Shihao, et al.
Published: (2024)
by: Xia, Shihao, et al.
Published: (2024)
Demystifying ChatGPT: How It Masters Genre Recognition
by: Raj, Subham, et al.
Published: (2025)
by: Raj, Subham, et al.
Published: (2025)
Can ChatGPT Really Understand Modern Chinese Poetry?
by: Wang, Shanshan, et al.
Published: (2026)
by: Wang, Shanshan, et al.
Published: (2026)
Evaluating the Performance of ChatGPT for Spam Email Detection
by: Si, Shijing, et al.
Published: (2024)
by: Si, Shijing, et al.
Published: (2024)
Developing ChatGPT for Biology and Medicine: A Complete Review of Biomedical Question Answering
by: Li, Qing, et al.
Published: (2024)
by: Li, Qing, et al.
Published: (2024)
Working Memory Capacity of ChatGPT: An Empirical Study
by: Gong, Dongyu, et al.
Published: (2023)
by: Gong, Dongyu, et al.
Published: (2023)
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023)
by: Aiyappa, Rachith, et al.
Published: (2023)
Comparative Evaluation of ChatGPT and DeepSeek Across Key NLP Tasks: Strengths, Weaknesses, and Domain-Specific Performance
by: Etaiwi, Wael, et al.
Published: (2025)
by: Etaiwi, Wael, et al.
Published: (2025)
Evaluating the Application of ChatGPT in Outpatient Triage Guidance: A Comparative Study
by: Liu, Dou, et al.
Published: (2024)
by: Liu, Dou, et al.
Published: (2024)
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Event-level Knowledge Editing
by: Peng, Hao, et al.
Published: (2024)
by: Peng, Hao, et al.
Published: (2024)
Evaluating ChatGPT as a Recommender System: A Rigorous Approach
by: Di Palma, Dario, et al.
Published: (2023)
by: Di Palma, Dario, et al.
Published: (2023)
ChatGPT Alternative Solutions: Large Language Models Survey
by: Alipour, Hanieh, et al.
Published: (2024)
by: Alipour, Hanieh, et al.
Published: (2024)
Experimental evidence of progressive ChatGPT models self-convergence
by: Xylogiannopoulos, Konstantinos F., et al.
Published: (2026)
by: Xylogiannopoulos, Konstantinos F., et al.
Published: (2026)
Similar Items
-
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2023) -
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
by: Tu, Shangqing, et al.
Published: (2024) -
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023) -
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023) -
Primacy Effect of ChatGPT
by: Wang, Yiwei, et al.
Published: (2023)