HumanEval on Latest GPT Models -- 2024
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Daniel, Murr, Lincoln |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reactor Mk.1 performances: MMLU, HumanEval and BBH test results
von: Dunham, TJ, et al.
Veröffentlicht: (2024)
von: Dunham, TJ, et al.
Veröffentlicht: (2024)
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2024)
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2024)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
IITK at SemEval-2024 Task 4: Hierarchical Embeddings for Detection of Persuasion Techniques in Memes
von: Chikoti, Shreenaga, et al.
Veröffentlicht: (2024)
von: Chikoti, Shreenaga, et al.
Veröffentlicht: (2024)
SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense
von: Jiang, Yifan, et al.
Veröffentlicht: (2024)
von: Jiang, Yifan, et al.
Veröffentlicht: (2024)
AIMA at SemEval-2024 Task 3: Simple Yet Powerful Emotion Cause Pair Analysis
von: Kure, Alireza Ghahramani, et al.
Veröffentlicht: (2025)
von: Kure, Alireza Ghahramani, et al.
Veröffentlicht: (2025)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts
von: Basak, Udvas, et al.
Veröffentlicht: (2024)
von: Basak, Udvas, et al.
Veröffentlicht: (2024)
ArabianGPT: Native Arabic GPT-based Large Language Model
von: Koubaa, Anis, et al.
Veröffentlicht: (2024)
von: Koubaa, Anis, et al.
Veröffentlicht: (2024)
ScholarEval: Research Idea Evaluation Grounded in Literature
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2025)
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2025)
UMBCLU at SemEval-2024 Task 1A and 1C: Semantic Textual Relatedness with and without machine translation
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2024)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2024)
TLDR at SemEval-2024 Task 2: T5-generated clinical-Language summaries for DeBERTa Report Analysis
von: Das, Spandan, et al.
Veröffentlicht: (2024)
von: Das, Spandan, et al.
Veröffentlicht: (2024)
AIMA at SemEval-2024 Task 10: History-Based Emotion Recognition in Hindi-English Code-Mixed Conversations
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
Measuring all the noises of LLM Evals
von: Wang, Sida
Veröffentlicht: (2025)
von: Wang, Sida
Veröffentlicht: (2025)
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials
von: Mandal, Shreyasi, et al.
Veröffentlicht: (2024)
von: Mandal, Shreyasi, et al.
Veröffentlicht: (2024)
HU at SemEval-2024 Task 8A: Can Contrastive Learning Learn Embeddings to Detect Machine-Generated Text?
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2024)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2024)
RKadiyala at SemEval-2024 Task 8: Black-Box Word-Level Text Boundary Detection in Partially Machine Generated Texts
von: Kadiyala, Ram Mohan Rao
Veröffentlicht: (2024)
von: Kadiyala, Ram Mohan Rao
Veröffentlicht: (2024)
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
von: Patel, Shubham, et al.
Veröffentlicht: (2024)
von: Patel, Shubham, et al.
Veröffentlicht: (2024)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Universal Neurons in GPT2 Language Models
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
CityGPT: Empowering Urban Spatial Cognition of Large Language Models
von: Feng, Jie, et al.
Veröffentlicht: (2024)
von: Feng, Jie, et al.
Veröffentlicht: (2024)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
AcademicEval: Live Long-Context LLM Benchmark
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
von: Tang, Zihao, et al.
Veröffentlicht: (2024)
von: Tang, Zihao, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
von: Mishra, Anurag
Veröffentlicht: (2025)
von: Mishra, Anurag
Veröffentlicht: (2025)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
von: Boyeau, Pierre, et al.
Veröffentlicht: (2024)
von: Boyeau, Pierre, et al.
Veröffentlicht: (2024)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
von: Aravapalli, Akhilesh, et al.
Veröffentlicht: (2024)
von: Aravapalli, Akhilesh, et al.
Veröffentlicht: (2024)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
von: Zain, Noor Ul, et al.
Veröffentlicht: (2025)
von: Zain, Noor Ul, et al.
Veröffentlicht: (2025)
Fairness of ChatGPT
von: Li, Yunqi, et al.
Veröffentlicht: (2023)
von: Li, Yunqi, et al.
Veröffentlicht: (2023)
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models
von: Lv, Yaojia, et al.
Veröffentlicht: (2024)
von: Lv, Yaojia, et al.
Veröffentlicht: (2024)
Using GPT Models for Qualitative and Quantitative News Analytics in the 2024 US Presidental Election Process
von: Pavlyshenko, Bohdan M.
Veröffentlicht: (2024)
von: Pavlyshenko, Bohdan M.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reactor Mk.1 performances: MMLU, HumanEval and BBH test results
von: Dunham, TJ, et al.
Veröffentlicht: (2024) -
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2024) -
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023) -
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023) -
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)