Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Qintong, Cui, Leyang, Kong, Lingpeng, Bi, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
von: Li, Qintong, et al.
Veröffentlicht: (2024)
von: Li, Qintong, et al.
Veröffentlicht: (2024)
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
MAGE: Machine-generated Text Detection in the Wild
von: Li, Yafu, et al.
Veröffentlicht: (2023)
von: Li, Yafu, et al.
Veröffentlicht: (2023)
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Developing and Utilizing a Large-Scale Cantonese Dataset for Multi-Tasking in Large Language Models
von: Jiang, Jiyue, et al.
Veröffentlicht: (2025)
von: Jiang, Jiyue, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models on Multiple Tasks in Bioinformatics NLP with Prompting
von: Jiang, Jiyue, et al.
Veröffentlicht: (2025)
von: Jiang, Jiyue, et al.
Veröffentlicht: (2025)
Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
von: Shi, Shuming, et al.
Veröffentlicht: (2024)
von: Shi, Shuming, et al.
Veröffentlicht: (2024)
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
von: Prange, Jakob, et al.
Veröffentlicht: (2021)
von: Prange, Jakob, et al.
Veröffentlicht: (2021)
DROJ: A Prompt-Driven Attack against Large Language Models
von: Hu, Leyang, et al.
Veröffentlicht: (2024)
von: Hu, Leyang, et al.
Veröffentlicht: (2024)
Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration
von: Li, Qintong, et al.
Veröffentlicht: (2024)
von: Li, Qintong, et al.
Veröffentlicht: (2024)
Training-Free Long-Context Scaling of Large Language Models
von: An, Chenxin, et al.
Veröffentlicht: (2024)
von: An, Chenxin, et al.
Veröffentlicht: (2024)
D-NLP at SemEval-2024 Task 2: Evaluating Clinical Inference Capabilities of Large Language Models
von: Altinok, Duygu
Veröffentlicht: (2024)
von: Altinok, Duygu
Veröffentlicht: (2024)
Dream 7B: Diffusion Large Language Models
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
Exploring Large Language Models in Healthcare: Insights into Corpora Sources, Customization Strategies, and Evaluation Metrics
von: Yang, Shuqi, et al.
Veröffentlicht: (2025)
von: Yang, Shuqi, et al.
Veröffentlicht: (2025)
Proxy Compression for Language Modeling
von: Zheng, Lin, et al.
Veröffentlicht: (2026)
von: Zheng, Lin, et al.
Veröffentlicht: (2026)
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
von: Gu, Yiyang, et al.
Veröffentlicht: (2026)
von: Gu, Yiyang, et al.
Veröffentlicht: (2026)
Neuro-Symbolic Integration Brings Causal and Reliable Reasoning Proofs
von: Yang, Sen, et al.
Veröffentlicht: (2023)
von: Yang, Sen, et al.
Veröffentlicht: (2023)
Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts
von: Gao, Fan, et al.
Veröffentlicht: (2023)
von: Gao, Fan, et al.
Veröffentlicht: (2023)
Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis
von: Zhu, Wenhao, et al.
Veröffentlicht: (2023)
von: Zhu, Wenhao, et al.
Veröffentlicht: (2023)
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
von: Aung, Thura, et al.
Veröffentlicht: (2026)
von: Aung, Thura, et al.
Veröffentlicht: (2026)
ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models
von: Han, Simeng, et al.
Veröffentlicht: (2025)
von: Han, Simeng, et al.
Veröffentlicht: (2025)
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
von: Gao, Jiahui, et al.
Veröffentlicht: (2024)
von: Gao, Jiahui, et al.
Veröffentlicht: (2024)
Scaling Diffusion Language Models via Adaptation from Autoregressive Models
von: Gong, Shansan, et al.
Veröffentlicht: (2024)
von: Gong, Shansan, et al.
Veröffentlicht: (2024)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
von: Jiang, Jiyue, et al.
Veröffentlicht: (2024)
von: Jiang, Jiyue, et al.
Veröffentlicht: (2024)
Knowledge Verification to Nip Hallucination in the Bud
von: Wan, Fanqi, et al.
Veröffentlicht: (2024)
von: Wan, Fanqi, et al.
Veröffentlicht: (2024)
xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
von: Yu, Qingchen, et al.
Veröffentlicht: (2024)
von: Yu, Qingchen, et al.
Veröffentlicht: (2024)
Active Learning for NLP with Large Language Models
von: Wang, Xuesong
Veröffentlicht: (2024)
von: Wang, Xuesong
Veröffentlicht: (2024)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size Canvas
von: Wu, Zirui, et al.
Veröffentlicht: (2026)
von: Wu, Zirui, et al.
Veröffentlicht: (2026)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
von: Li, Yafu, et al.
Veröffentlicht: (2024)
von: Li, Yafu, et al.
Veröffentlicht: (2024)
NLP Datasets for Idiom and Figurative Language Tasks
von: Matheny, Blake, et al.
Veröffentlicht: (2025)
von: Matheny, Blake, et al.
Veröffentlicht: (2025)
A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
Evaluating the Reliability of Self-Explanations in Large Language Models
von: Randl, Korbinian, et al.
Veröffentlicht: (2024)
von: Randl, Korbinian, et al.
Veröffentlicht: (2024)
Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
NLP needs Diversity outside of 'Diversity'
von: Tint, Joshua
Veröffentlicht: (2026)
von: Tint, Joshua
Veröffentlicht: (2026)
Activation-Guided Consensus Merging for Large Language Models
von: Yao, Yuxuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
von: Li, Qintong, et al.
Veröffentlicht: (2024) -
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025) -
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024) -
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
von: Zhang, Yue, et al.
Veröffentlicht: (2023) -
MAGE: Machine-generated Text Detection in the Wild
von: Li, Yafu, et al.
Veröffentlicht: (2023)