IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tran, Khanh-Tung, O'Sullivan, Barry, Nguyen, Hoang D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UCCIX: Irish-eXcellence Large Language Model
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2024)
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2024)
Reasoning Transfer for an Extremely Low-Resource and Endangered Language: Bridging Languages Through Sample-Efficient Language Understanding
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
von: McGiff, Josh, et al.
Veröffentlicht: (2025)
von: McGiff, Josh, et al.
Veröffentlicht: (2025)
Qomhra: A Bilingual Irish and English Large Language Model
von: McInerney, Joseph, et al.
Veröffentlicht: (2025)
von: McInerney, Joseph, et al.
Veröffentlicht: (2025)
Multi-Agent Collaboration Mechanisms: A Survey of LLMs
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
Questionnaire meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
von: Nguyen, Duc-Hai, et al.
Veröffentlicht: (2025)
von: Nguyen, Duc-Hai, et al.
Veröffentlicht: (2025)
Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
von: Loc, Ngoc Phan Phuoc, et al.
Veröffentlicht: (2026)
von: Loc, Ngoc Phan Phuoc, et al.
Veröffentlicht: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
OpenSIR: Open-Ended Self-Improving Reasoner
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2025)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2025)
ViFactCheck: A New Benchmark Dataset and Methods for Multi-domain News Fact-Checking in Vietnamese
von: Hoa, Tran Thai, et al.
Veröffentlicht: (2024)
von: Hoa, Tran Thai, et al.
Veröffentlicht: (2024)
MASIVE: Open-Ended Affective State Identification in English and Spanish
von: Deas, Nicholas, et al.
Veröffentlicht: (2024)
von: Deas, Nicholas, et al.
Veröffentlicht: (2024)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
Reverse-Engineered Reasoning for Open-Ended Generation
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
Scaling Open-Ended Reasoning to Predict the Future
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Human Evaluation of English--Irish Transformer-Based NMT
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
OWLViz: An Open-World Benchmark for Visual Question Answering
von: Nguyen, Thuy, et al.
Veröffentlicht: (2025)
von: Nguyen, Thuy, et al.
Veröffentlicht: (2025)
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
von: Bhatti, Hunzalah Hassan, et al.
Veröffentlicht: (2025)
von: Bhatti, Hunzalah Hassan, et al.
Veröffentlicht: (2025)
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
BlasBench: An Open Benchmark for Irish Speech Recognition
von: Raj, Jyoutir, et al.
Veröffentlicht: (2026)
von: Raj, Jyoutir, et al.
Veröffentlicht: (2026)
Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations
von: Rinki, Mamnuya, et al.
Veröffentlicht: (2025)
von: Rinki, Mamnuya, et al.
Veröffentlicht: (2025)
Semantic Agreement Enables Efficient Open-Ended LLM Cascades
von: Soiffer, Duncan, et al.
Veröffentlicht: (2025)
von: Soiffer, Duncan, et al.
Veröffentlicht: (2025)
Bias Association Discovery Framework for Open-Ended LLM Generations
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs
von: Ye, Yangfan, et al.
Veröffentlicht: (2026)
von: Ye, Yangfan, et al.
Veröffentlicht: (2026)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
An Answer is just the Start: Related Insight Generation for Open-Ended Document-Grounded QA
von: Sharma, Saransh, et al.
Veröffentlicht: (2026)
von: Sharma, Saransh, et al.
Veröffentlicht: (2026)
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
CheckEmbed: Effective Verification of LLM Solutions to Open-Ended Tasks
von: Besta, Maciej, et al.
Veröffentlicht: (2024)
von: Besta, Maciej, et al.
Veröffentlicht: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
von: Li, Yu, et al.
Veröffentlicht: (2024)
von: Li, Yu, et al.
Veröffentlicht: (2024)
End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models
von: Zhang, Yunhao, et al.
Veröffentlicht: (2024)
von: Zhang, Yunhao, et al.
Veröffentlicht: (2024)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
von: Pham, Loc, et al.
Veröffentlicht: (2025)
von: Pham, Loc, et al.
Veröffentlicht: (2025)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
von: Amirizaniani, Maryam, et al.
Veröffentlicht: (2024)
von: Amirizaniani, Maryam, et al.
Veröffentlicht: (2024)
VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding
von: Do, Phong Nguyen-Thuan, et al.
Veröffentlicht: (2024)
von: Do, Phong Nguyen-Thuan, et al.
Veröffentlicht: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
UCCIX: Irish-eXcellence Large Language Model
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2024) -
Reasoning Transfer for an Extremely Low-Resource and Endangered Language: Bridging Languages Through Sample-Efficient Language Understanding
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025) -
Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
von: McGiff, Josh, et al.
Veröffentlicht: (2025) -
Qomhra: A Bilingual Irish and English Large Language Model
von: McInerney, Joseph, et al.
Veröffentlicht: (2025) -
Multi-Agent Collaboration Mechanisms: A Survey of LLMs
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)