PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Niazi, Ruhallah, Ghorbanpour, Faeze, Fraser, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
Are BabyLMs Second Language Learners?
von: Edman, Lukas, et al.
Veröffentlicht: (2024)
von: Edman, Lukas, et al.
Veröffentlicht: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
Differentiating Emigration from Return Migration of Scholars Using Name-Based Nationality Detection Models
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
LitLLMs, LLMs for Literature Review: Are we there yet?
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
von: Li, Sihang, et al.
Veröffentlicht: (2024)
von: Li, Sihang, et al.
Veröffentlicht: (2024)
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language
von: Jamali, Naghmeh, et al.
Veröffentlicht: (2025)
von: Jamali, Naghmeh, et al.
Veröffentlicht: (2025)
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
von: Monazzah, Erfan Moosavi, et al.
Veröffentlicht: (2025)
von: Monazzah, Erfan Moosavi, et al.
Veröffentlicht: (2025)
MAS-LitEval : Multi-Agent System for Literary Translation Quality Assessment
von: Kim, Junghwan, et al.
Veröffentlicht: (2025)
von: Kim, Junghwan, et al.
Veröffentlicht: (2025)
LitSearch: A Retrieval Benchmark for Scientific Literature Search
von: Ajith, Anirudh, et al.
Veröffentlicht: (2024)
von: Ajith, Anirudh, et al.
Veröffentlicht: (2024)
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
von: Fein, Daniel, et al.
Veröffentlicht: (2025)
von: Fein, Daniel, et al.
Veröffentlicht: (2025)
Conversational Exploration of Literature Landscape with LitChat
von: Huang, Mingyu, et al.
Veröffentlicht: (2025)
von: Huang, Mingyu, et al.
Veröffentlicht: (2025)
TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
von: Iranmanesh, Reihaneh, et al.
Veröffentlicht: (2026)
von: Iranmanesh, Reihaneh, et al.
Veröffentlicht: (2026)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
FineSurE: Fine-grained Summarization Evaluation using LLMs
von: Song, Hwanjun, et al.
Veröffentlicht: (2024)
von: Song, Hwanjun, et al.
Veröffentlicht: (2024)
HKCanto-Eval: A Benchmark for Evaluating Cantonese Language Understanding and Cultural Comprehension in LLMs
von: Cheng, Tsz Chung, et al.
Veröffentlicht: (2025)
von: Cheng, Tsz Chung, et al.
Veröffentlicht: (2025)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
LitLLM: A Toolkit for Scientific Literature Review
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
von: Huang, Shulin, et al.
Veröffentlicht: (2023)
von: Huang, Shulin, et al.
Veröffentlicht: (2023)
PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2026)
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2026)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
von: Lee, Yuho, et al.
Veröffentlicht: (2024)
von: Lee, Yuho, et al.
Veröffentlicht: (2024)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
LitVISTA: A Benchmark for Narrative Orchestration in Literary Text
von: Lu, Mingzhe, et al.
Veröffentlicht: (2026)
von: Lu, Mingzhe, et al.
Veröffentlicht: (2026)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
von: Du, Junjia, et al.
Veröffentlicht: (2025)
von: Du, Junjia, et al.
Veröffentlicht: (2025)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
StackEval: Benchmarking LLMs in Coding Assistance
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question Answering
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback
von: Lai, Wen, et al.
Veröffentlicht: (2024)
von: Lai, Wen, et al.
Veröffentlicht: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023)
von: Shu, Lei, et al.
Veröffentlicht: (2023)
Evaluating the Creativity of LLMs in Persian Literary Text Generation
von: Tourajmehr, Armin, et al.
Veröffentlicht: (2025)
von: Tourajmehr, Armin, et al.
Veröffentlicht: (2025)
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
von: Mendonça, John, et al.
Veröffentlicht: (2024)
von: Mendonça, John, et al.
Veröffentlicht: (2024)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
MiLiC-Eval: Benchmarking Multilingual LLMs for China's Minority Languages
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
CUTE: Measuring LLMs' Understanding of Their Tokens
von: Edman, Lukas, et al.
Veröffentlicht: (2024)
von: Edman, Lukas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025) -
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025) -
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025) -
Are BabyLMs Second Language Learners?
von: Edman, Lukas, et al.
Veröffentlicht: (2024) -
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)