TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Iranmanesh, Reihaneh, Davoudi, Saeedeh, Abrishamchian, Pasha, Frieder, Ophir, Goharian, Nazli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis
von: Davoudi, Saeedeh, et al.
Veröffentlicht: (2026)
von: Davoudi, Saeedeh, et al.
Veröffentlicht: (2026)
Genetic Approach to Mitigate Hallucination in Generative IR
von: Kulkarni, Hrishikesh, et al.
Veröffentlicht: (2024)
von: Kulkarni, Hrishikesh, et al.
Veröffentlicht: (2024)
Intercept Cancer: Cancer Pre-Screening with Large Scale Healthcare Foundation Models
von: Sun, Liwen, et al.
Veröffentlicht: (2025)
von: Sun, Liwen, et al.
Veröffentlicht: (2025)
Learning to Rank Salient Content for Query-focused Summarization
von: Sotudeh, Sajad, et al.
Veröffentlicht: (2024)
von: Sotudeh, Sajad, et al.
Veröffentlicht: (2024)
ASAG2024: A Combined Benchmark for Short Answer Grading
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024)
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
LexBoost: Improving Lexical Document Retrieval with Nearest Neighbors
von: Kulkarni, Hrishikesh, et al.
Veröffentlicht: (2024)
von: Kulkarni, Hrishikesh, et al.
Veröffentlicht: (2024)
Question-Answer Extraction from Scientific Articles Using Knowledge Graphs and Large Language Models
von: Azarbonyad, Hosein, et al.
Veröffentlicht: (2025)
von: Azarbonyad, Hosein, et al.
Veröffentlicht: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2026)
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2026)
Persian Slang Text Conversion to Formal and Deep Learning of Persian Short Texts on Social Media for Sentiment Classification
von: Khazeni, Mohsen, et al.
Veröffentlicht: (2024)
von: Khazeni, Mohsen, et al.
Veröffentlicht: (2024)
Benchmarking Uncertainty Calibration in Large Language Model Long-Form Question Answering
von: Müller, Philip, et al.
Veröffentlicht: (2026)
von: Müller, Philip, et al.
Veröffentlicht: (2026)
Generating Text from Uniform Meaning Representation
von: Markle, Emma, et al.
Veröffentlicht: (2025)
von: Markle, Emma, et al.
Veröffentlicht: (2025)
A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2025)
von: Rogoz, Ana-Cristina, et al.
Veröffentlicht: (2025)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
FaMTEB: Massive Text Embedding Benchmark in Persian Language
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
Are Large Language Models Good Temporal Graph Learners?
von: Huang, Shenyang, et al.
Veröffentlicht: (2025)
von: Huang, Shenyang, et al.
Veröffentlicht: (2025)
Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning
von: Gado, Elena Grazia, et al.
Veröffentlicht: (2024)
von: Gado, Elena Grazia, et al.
Veröffentlicht: (2024)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
von: Wang, Song, et al.
Veröffentlicht: (2024)
von: Wang, Song, et al.
Veröffentlicht: (2024)
Large Language Models for Mathematicians
von: Frieder, Simon, et al.
Veröffentlicht: (2023)
von: Frieder, Simon, et al.
Veröffentlicht: (2023)
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
von: Farsi, Farhan, et al.
Veröffentlicht: (2025)
von: Farsi, Farhan, et al.
Veröffentlicht: (2025)
Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition
von: Uddin, Mohammed Mudassir, et al.
Veröffentlicht: (2026)
von: Uddin, Mohammed Mudassir, et al.
Veröffentlicht: (2026)
CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answering
von: Liu, Yu, et al.
Veröffentlicht: (2026)
von: Liu, Yu, et al.
Veröffentlicht: (2026)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
von: Seegmiller, Parker, et al.
Veröffentlicht: (2024)
von: Seegmiller, Parker, et al.
Veröffentlicht: (2024)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
von: Bonthu, Sridevi, et al.
Veröffentlicht: (2025)
von: Bonthu, Sridevi, et al.
Veröffentlicht: (2025)
Language Models Entangle Language and Culture
von: Jain, Shourya, et al.
Veröffentlicht: (2026)
von: Jain, Shourya, et al.
Veröffentlicht: (2026)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
von: Saadat, Asir, et al.
Veröffentlicht: (2024)
von: Saadat, Asir, et al.
Veröffentlicht: (2024)
Neural Isomorphic Fields: A Transformer-based Algebraic Numerical Embedding
von: Sadeghi, Hamidreza, et al.
Veröffentlicht: (2026)
von: Sadeghi, Hamidreza, et al.
Veröffentlicht: (2026)
On the Robustness of Answer Formats in Medical Reasoning Models
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2025)
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2025)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
von: Haussmann, Aden
Veröffentlicht: (2025)
von: Haussmann, Aden
Veröffentlicht: (2025)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
von: Park, Jean, et al.
Veröffentlicht: (2024)
von: Park, Jean, et al.
Veröffentlicht: (2024)
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
Calibrated Large Language Models for Binary Question Answering
von: Giovannotti, Patrizio, et al.
Veröffentlicht: (2024)
von: Giovannotti, Patrizio, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis
von: Davoudi, Saeedeh, et al.
Veröffentlicht: (2026) -
Genetic Approach to Mitigate Hallucination in Generative IR
von: Kulkarni, Hrishikesh, et al.
Veröffentlicht: (2024) -
Intercept Cancer: Cancer Pre-Screening with Large Scale Healthcare Foundation Models
von: Sun, Liwen, et al.
Veröffentlicht: (2025) -
Learning to Rank Salient Content for Query-focused Summarization
von: Sotudeh, Sajad, et al.
Veröffentlicht: (2024) -
ASAG2024: A Combined Benchmark for Short Answer Grading
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024)