Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mercorio, Fabio, Mezzanzanica, Mario, Potertì, Daniele, Serino, Antonio, Seveso, Andrea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Designing Role Vectors to Improve LLM Inference Behaviour
von: Potertì, Daniele, et al.
Veröffentlicht: (2025)
von: Potertì, Daniele, et al.
Veröffentlicht: (2025)
Towards the Terminator Economy: Assessing Job Exposure to AI through LLMs
von: Colombo, Emilio, et al.
Veröffentlicht: (2024)
von: Colombo, Emilio, et al.
Veröffentlicht: (2024)
XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
von: Cambria, Erik, et al.
Veröffentlicht: (2024)
von: Cambria, Erik, et al.
Veröffentlicht: (2024)
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
von: Sassella, Andrea, et al.
Veröffentlicht: (2026)
von: Sassella, Andrea, et al.
Veröffentlicht: (2026)
GemMaroc: Unlocking Darija Proficiency in LLMs with Minimal Data
von: Skiredj, Abderrahman, et al.
Veröffentlicht: (2025)
von: Skiredj, Abderrahman, et al.
Veröffentlicht: (2025)
How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian
von: Pedrotti, Andrea, et al.
Veröffentlicht: (2025)
von: Pedrotti, Andrea, et al.
Veröffentlicht: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
von: Lunardi, Riccardo, et al.
Veröffentlicht: (2025)
von: Lunardi, Riccardo, et al.
Veröffentlicht: (2025)
Harnessing LLMs for Educational Content-Driven Italian Crossword Generation
von: Zeinalipour, Kamyar, et al.
Veröffentlicht: (2024)
von: Zeinalipour, Kamyar, et al.
Veröffentlicht: (2024)
Evaluating LLMs on Entity Disambiguation in Tables
von: Belotti, Federico, et al.
Veröffentlicht: (2024)
von: Belotti, Federico, et al.
Veröffentlicht: (2024)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency
von: Shahriar, Sakib, et al.
Veröffentlicht: (2024)
von: Shahriar, Sakib, et al.
Veröffentlicht: (2024)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
von: Agarwal, Parth, et al.
Veröffentlicht: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2024)
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2024)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
von: Fabbri, Alexander R., et al.
Veröffentlicht: (2025)
von: Fabbri, Alexander R., et al.
Veröffentlicht: (2025)
\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments
von: Ghosh, Debanjan, et al.
Veröffentlicht: (2024)
von: Ghosh, Debanjan, et al.
Veröffentlicht: (2024)
Aligning Sentence Simplification with ESL Learner's Proficiency for Language Acquisition
von: Li, Guanlin, et al.
Veröffentlicht: (2025)
von: Li, Guanlin, et al.
Veröffentlicht: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2024)
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2024)
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
von: Guo, Qianhong, et al.
Veröffentlicht: (2025)
von: Guo, Qianhong, et al.
Veröffentlicht: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
Can Large Language Models Automatically Score Proficiency of Written Essays?
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
Classifying German Language Proficiency Levels Using Large Language Models
von: Ahlers, Elias-Leander, et al.
Veröffentlicht: (2025)
von: Ahlers, Elias-Leander, et al.
Veröffentlicht: (2025)
Protoknowledge Shapes Behaviour of LLMs in Downstream Tasks: Memorization and Generalization with Knowledge Graphs
von: Ranaldi, Federico, et al.
Veröffentlicht: (2025)
von: Ranaldi, Federico, et al.
Veröffentlicht: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
TimeSense:Making Large Language Models Proficient in Time-Series Analysis
von: Zhang, Zhirui, et al.
Veröffentlicht: (2025)
von: Zhang, Zhirui, et al.
Veröffentlicht: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Designing Role Vectors to Improve LLM Inference Behaviour
von: Potertì, Daniele, et al.
Veröffentlicht: (2025) -
Towards the Terminator Economy: Assessing Job Exposure to AI through LLMs
von: Colombo, Emilio, et al.
Veröffentlicht: (2024) -
XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
von: Cambria, Erik, et al.
Veröffentlicht: (2024) -
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025) -
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
von: Sassella, Andrea, et al.
Veröffentlicht: (2026)