Salvato in:
| Autori principali: | Wu, Kevin, Wu, Eric, Zou, James |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2404.10198 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
di: Wu, Eric, et al.
Pubblicazione: (2024)
di: Wu, Eric, et al.
Pubblicazione: (2024)
TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy
di: McCammon, James
Pubblicazione: (2025)
di: McCammon, James
Pubblicazione: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
di: Wu, JiaRu, et al.
Pubblicazione: (2025)
di: Wu, JiaRu, et al.
Pubblicazione: (2025)
System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection
di: Wu, Binglin, et al.
Pubblicazione: (2025)
di: Wu, Binglin, et al.
Pubblicazione: (2025)
How well do LLMs cite relevant medical references? An evaluation framework and analyses
di: Wu, Kevin, et al.
Pubblicazione: (2024)
di: Wu, Kevin, et al.
Pubblicazione: (2024)
Tug-of-war between idioms' figurative and literal interpretations in LLMs
di: Oh, Soyoung, et al.
Pubblicazione: (2025)
di: Oh, Soyoung, et al.
Pubblicazione: (2025)
PatentEval: Understanding Errors in Patent Generation
di: Zuo, You, et al.
Pubblicazione: (2024)
di: Zuo, You, et al.
Pubblicazione: (2024)
Measuring all the noises of LLM Evals
di: Wang, Sida
Pubblicazione: (2025)
di: Wang, Sida
Pubblicazione: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
Data Compressibility Quantifies LLM Memorization
di: Huang, Yizhan, et al.
Pubblicazione: (2025)
di: Huang, Yizhan, et al.
Pubblicazione: (2025)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
di: D'Souza, Jennifer, et al.
Pubblicazione: (2025)
di: D'Souza, Jennifer, et al.
Pubblicazione: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
A Single Character can Make or Break Your LLM Evals
di: Su, Jingtong, et al.
Pubblicazione: (2025)
di: Su, Jingtong, et al.
Pubblicazione: (2025)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
di: Khan, Haidar, et al.
Pubblicazione: (2025)
di: Khan, Haidar, et al.
Pubblicazione: (2025)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
di: Zhao, Jiahao, et al.
Pubblicazione: (2025)
di: Zhao, Jiahao, et al.
Pubblicazione: (2025)
Science Across Languages: Assessing LLM Multilingual Translation of Scientific Papers
di: Kleidermacher, Hannah Calzi, et al.
Pubblicazione: (2025)
di: Kleidermacher, Hannah Calzi, et al.
Pubblicazione: (2025)
BriLLM: Brain-inspired Large Language Model
di: Zhao, Hai, et al.
Pubblicazione: (2025)
di: Zhao, Hai, et al.
Pubblicazione: (2025)
ReasonOps: Operator Segmentation for LLM Reasoning Traces
di: Lee, Daniel, et al.
Pubblicazione: (2026)
di: Lee, Daniel, et al.
Pubblicazione: (2026)
An evaluation of LLMs for political bias in Western media: Israel-Hamas and Ukraine-Russia wars
di: Chandra, Rohitash, et al.
Pubblicazione: (2026)
di: Chandra, Rohitash, et al.
Pubblicazione: (2026)
CausalEval: Towards Better Causal Reasoning in Language Models
di: Yu, Longxuan, et al.
Pubblicazione: (2024)
di: Yu, Longxuan, et al.
Pubblicazione: (2024)
AcademicEval: Live Long-Context LLM Benchmark
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
Disentangling Reasoning and Knowledge in Medical Large Language Models
di: Thapa, Rahul, et al.
Pubblicazione: (2025)
di: Thapa, Rahul, et al.
Pubblicazione: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
di: Nguyen, Trong-Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Trong-Hieu, et al.
Pubblicazione: (2024)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
di: Alyahya, Hisham A., et al.
Pubblicazione: (2025)
di: Alyahya, Hisham A., et al.
Pubblicazione: (2025)
Universal Legal Article Prediction via Tight Collaboration between Supervised Classification Model and LLM
di: Chi, Xiao, et al.
Pubblicazione: (2025)
di: Chi, Xiao, et al.
Pubblicazione: (2025)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
di: Zhang, Ziqian, et al.
Pubblicazione: (2026)
di: Zhang, Ziqian, et al.
Pubblicazione: (2026)
UoR-NCL at SemEval-2025 Task 1: Using Generative LLMs and CLIP Models for Multilingual Multimodal Idiomaticity Representation
di: Markchom, Thanet, et al.
Pubblicazione: (2025)
di: Markchom, Thanet, et al.
Pubblicazione: (2025)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
di: Huang, Zhongzhan, et al.
Pubblicazione: (2025)
di: Huang, Zhongzhan, et al.
Pubblicazione: (2025)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
di: Yu, Linhao, et al.
Pubblicazione: (2024)
di: Yu, Linhao, et al.
Pubblicazione: (2024)
Quantifying Geospatial in the Common Crawl Corpus
di: Ilyankou, Ilya, et al.
Pubblicazione: (2024)
di: Ilyankou, Ilya, et al.
Pubblicazione: (2024)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
di: Vemula, Saketh Reddy, et al.
Pubblicazione: (2025)
di: Vemula, Saketh Reddy, et al.
Pubblicazione: (2025)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
di: Zhao, Jingqian, et al.
Pubblicazione: (2025)
di: Zhao, Jingqian, et al.
Pubblicazione: (2025)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
Unlocking LLM Creativity in Science through Analogical Reasoning
di: Shen, Andrew, et al.
Pubblicazione: (2026)
di: Shen, Andrew, et al.
Pubblicazione: (2026)
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
di: Ma, Zi-Ao, et al.
Pubblicazione: (2025)
di: Ma, Zi-Ao, et al.
Pubblicazione: (2025)
SHROOM-INDElab at SemEval-2024 Task 6: Zero- and Few-Shot LLM-Based Classification for Hallucination Detection
di: Allen, Bradley P., et al.
Pubblicazione: (2024)
di: Allen, Bradley P., et al.
Pubblicazione: (2024)
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
di: Bianchi, Federico, et al.
Pubblicazione: (2025)
di: Bianchi, Federico, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
di: Wu, Eric, et al.
Pubblicazione: (2024) -
TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy
di: McCammon, James
Pubblicazione: (2025) -
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
di: Wu, JiaRu, et al.
Pubblicazione: (2025) -
System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection
di: Wu, Binglin, et al.
Pubblicazione: (2025) -
How well do LLMs cite relevant medical references? An evaluation framework and analyses
di: Wu, Kevin, et al.
Pubblicazione: (2024)