Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Mercorio, Fabio, Mezzanzanica, Mario, Potertì, Daniele, Serino, Antonio, Seveso, Andrea |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Designing Role Vectors to Improve LLM Inference Behaviour
di: Potertì, Daniele, et al.
Pubblicazione: (2025)
di: Potertì, Daniele, et al.
Pubblicazione: (2025)
Towards the Terminator Economy: Assessing Job Exposure to AI through LLMs
di: Colombo, Emilio, et al.
Pubblicazione: (2024)
di: Colombo, Emilio, et al.
Pubblicazione: (2024)
XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
di: Cambria, Erik, et al.
Pubblicazione: (2024)
di: Cambria, Erik, et al.
Pubblicazione: (2024)
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
di: Sassella, Andrea, et al.
Pubblicazione: (2026)
di: Sassella, Andrea, et al.
Pubblicazione: (2026)
GemMaroc: Unlocking Darija Proficiency in LLMs with Minimal Data
di: Skiredj, Abderrahman, et al.
Pubblicazione: (2025)
di: Skiredj, Abderrahman, et al.
Pubblicazione: (2025)
How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian
di: Pedrotti, Andrea, et al.
Pubblicazione: (2025)
di: Pedrotti, Andrea, et al.
Pubblicazione: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
di: Jiang, Botian, et al.
Pubblicazione: (2024)
di: Jiang, Botian, et al.
Pubblicazione: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
di: Wang, Yicheng, et al.
Pubblicazione: (2024)
di: Wang, Yicheng, et al.
Pubblicazione: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
di: Lunardi, Riccardo, et al.
Pubblicazione: (2025)
di: Lunardi, Riccardo, et al.
Pubblicazione: (2025)
Harnessing LLMs for Educational Content-Driven Italian Crossword Generation
di: Zeinalipour, Kamyar, et al.
Pubblicazione: (2024)
di: Zeinalipour, Kamyar, et al.
Pubblicazione: (2024)
Evaluating LLMs on Entity Disambiguation in Tables
di: Belotti, Federico, et al.
Pubblicazione: (2024)
di: Belotti, Federico, et al.
Pubblicazione: (2024)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
di: Yadav, Ankit, et al.
Pubblicazione: (2024)
di: Yadav, Ankit, et al.
Pubblicazione: (2024)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
di: Wang, Jun, et al.
Pubblicazione: (2025)
di: Wang, Jun, et al.
Pubblicazione: (2025)
Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency
di: Shahriar, Sakib, et al.
Pubblicazione: (2024)
di: Shahriar, Sakib, et al.
Pubblicazione: (2024)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
di: Basu, Kinjal, et al.
Pubblicazione: (2024)
di: Basu, Kinjal, et al.
Pubblicazione: (2024)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
di: Zhang, Mengyuan, et al.
Pubblicazione: (2024)
di: Zhang, Mengyuan, et al.
Pubblicazione: (2024)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
di: Truong, Kimberly Le, et al.
Pubblicazione: (2025)
di: Truong, Kimberly Le, et al.
Pubblicazione: (2025)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
di: Agarwal, Parth, et al.
Pubblicazione: (2025)
di: Agarwal, Parth, et al.
Pubblicazione: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
di: Xu, Wenda, et al.
Pubblicazione: (2025)
di: Xu, Wenda, et al.
Pubblicazione: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
di: Bannò, Stefano, et al.
Pubblicazione: (2025)
di: Bannò, Stefano, et al.
Pubblicazione: (2025)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
di: Wang, Wanying, et al.
Pubblicazione: (2024)
di: Wang, Wanying, et al.
Pubblicazione: (2024)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
di: Chernyshev, Konstantin, et al.
Pubblicazione: (2024)
di: Chernyshev, Konstantin, et al.
Pubblicazione: (2024)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
di: Fabbri, Alexander R., et al.
Pubblicazione: (2025)
di: Fabbri, Alexander R., et al.
Pubblicazione: (2025)
\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments
di: Ghosh, Debanjan, et al.
Pubblicazione: (2024)
di: Ghosh, Debanjan, et al.
Pubblicazione: (2024)
Aligning Sentence Simplification with ESL Learner's Proficiency for Language Acquisition
di: Li, Guanlin, et al.
Pubblicazione: (2025)
di: Li, Guanlin, et al.
Pubblicazione: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
di: Zeng, Zhongshen, et al.
Pubblicazione: (2024)
di: Zeng, Zhongshen, et al.
Pubblicazione: (2024)
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
di: Guo, Qianhong, et al.
Pubblicazione: (2025)
di: Guo, Qianhong, et al.
Pubblicazione: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
Can Large Language Models Automatically Score Proficiency of Written Essays?
di: Mansour, Watheq, et al.
Pubblicazione: (2024)
di: Mansour, Watheq, et al.
Pubblicazione: (2024)
Classifying German Language Proficiency Levels Using Large Language Models
di: Ahlers, Elias-Leander, et al.
Pubblicazione: (2025)
di: Ahlers, Elias-Leander, et al.
Pubblicazione: (2025)
Protoknowledge Shapes Behaviour of LLMs in Downstream Tasks: Memorization and Generalization with Knowledge Graphs
di: Ranaldi, Federico, et al.
Pubblicazione: (2025)
di: Ranaldi, Federico, et al.
Pubblicazione: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
di: Wei, Jianhui, et al.
Pubblicazione: (2025)
di: Wei, Jianhui, et al.
Pubblicazione: (2025)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
di: Sirdeshmukh, Ved, et al.
Pubblicazione: (2025)
di: Sirdeshmukh, Ved, et al.
Pubblicazione: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
TimeSense:Making Large Language Models Proficient in Time-Series Analysis
di: Zhang, Zhirui, et al.
Pubblicazione: (2025)
di: Zhang, Zhirui, et al.
Pubblicazione: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Designing Role Vectors to Improve LLM Inference Behaviour
di: Potertì, Daniele, et al.
Pubblicazione: (2025) -
Towards the Terminator Economy: Assessing Job Exposure to AI through LLMs
di: Colombo, Emilio, et al.
Pubblicazione: (2024) -
XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
di: Cambria, Erik, et al.
Pubblicazione: (2024) -
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025) -
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
di: Sassella, Andrea, et al.
Pubblicazione: (2026)