All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Testa, Davide, Bonetta, Giovanni, Bernardi, Raffaella, Bondielli, Alessandro, Lenci, Alessandro, Miaschi, Alessio, Passaro, Lucia, Magnini, Bernardo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
von: Capone, Luca, et al.
Veröffentlicht: (2025)
von: Capone, Luca, et al.
Veröffentlicht: (2025)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
von: Miliani, Martina, et al.
Veröffentlicht: (2025)
von: Miliani, Martina, et al.
Veröffentlicht: (2025)
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
An Experimental Comparison of the Most Popular Approaches to Fake News Detection
von: Dell'Oglio, Pietro, et al.
Veröffentlicht: (2026)
von: Dell'Oglio, Pietro, et al.
Veröffentlicht: (2026)
Enhancing Debunking Effectiveness through LLM-based Personality Adaptation
von: Dell'Oglio, Pietro, et al.
Veröffentlicht: (2026)
von: Dell'Oglio, Pietro, et al.
Veröffentlicht: (2026)
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
von: Alzetta, Chiara, et al.
Veröffentlicht: (2025)
von: Alzetta, Chiara, et al.
Veröffentlicht: (2025)
Doing Things with Words: Rethinking Theory of Mind Simulation in Large Language Models
von: Lombardi, Agnese, et al.
Veröffentlicht: (2025)
von: Lombardi, Agnese, et al.
Veröffentlicht: (2025)
The quasi-semantic competence of LLMs: a case study on the part-whole relation
von: Proietti, Mattia, et al.
Veröffentlicht: (2025)
von: Proietti, Mattia, et al.
Veröffentlicht: (2025)
Evaluating Task-Oriented Dialogue Consistency through Constraint Satisfaction
von: Labruna, Tiziano, et al.
Veröffentlicht: (2024)
von: Labruna, Tiziano, et al.
Veröffentlicht: (2024)
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
von: Mazzaccara, Davide, et al.
Veröffentlicht: (2024)
von: Mazzaccara, Davide, et al.
Veröffentlicht: (2024)
Linguistic Knowledge Can Enhance Encoder-Decoder Models (If You Let It)
von: Miaschi, Alessio, et al.
Veröffentlicht: (2024)
von: Miaschi, Alessio, et al.
Veröffentlicht: (2024)
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2024)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2024)
Composing or Not Composing? Towards Distributional Construction Grammars
von: Blache, Philippe, et al.
Veröffentlicht: (2024)
von: Blache, Philippe, et al.
Veröffentlicht: (2024)
Converting Annotated Clinical Cases into Structured Case Report Forms
von: Ferrazzi, Pietro, et al.
Veröffentlicht: (2025)
von: Ferrazzi, Pietro, et al.
Veröffentlicht: (2025)
BAMBI: Developing Baby Language Models for Italian
von: Suozzi, Alice, et al.
Veröffentlicht: (2025)
von: Suozzi, Alice, et al.
Veröffentlicht: (2025)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
Small LLMs for Medical NLP: a Systematic Analysis of Few-Shot, Constraint Decoding, Fine-Tuning and Continual Pre-Training in Italian
von: Ferrazzi, Pietro, et al.
Veröffentlicht: (2026)
von: Ferrazzi, Pietro, et al.
Veröffentlicht: (2026)
Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
von: Pedrotti, Andrea, et al.
Veröffentlicht: (2025)
von: Pedrotti, Andrea, et al.
Veröffentlicht: (2025)
Evalita-LLM: Benchmarking Large Language Models on Italian
von: Magnini, Bernardo, et al.
Veröffentlicht: (2025)
von: Magnini, Bernardo, et al.
Veröffentlicht: (2025)
Probing for the Usage of Grammatical Number
von: Lasri, Karim, et al.
Veröffentlicht: (2022)
von: Lasri, Karim, et al.
Veröffentlicht: (2022)
Linguistic Profiling of a Neural Language Model
von: Miaschi, Alessio, et al.
Veröffentlicht: (2020)
von: Miaschi, Alessio, et al.
Veröffentlicht: (2020)
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models
von: Kauf, Carina, et al.
Veröffentlicht: (2024)
von: Kauf, Carina, et al.
Veröffentlicht: (2024)
Leveraging Encoder-only Large Language Models for Mobile App Review Feature Extraction
von: Motger, Quim, et al.
Veröffentlicht: (2024)
von: Motger, Quim, et al.
Veröffentlicht: (2024)
Toward Automatic Filling of Case Report Forms: A Case Study on Data from an Italian Emergency Department
von: Kaczmarek, Gabriela Anna, et al.
Veröffentlicht: (2026)
von: Kaczmarek, Gabriela Anna, et al.
Veröffentlicht: (2026)
Collocation in the Mind: Investigating Collocational Priming in Second Language Speakers of Italian
von: Irene Fioravanti, et al.
Veröffentlicht: (2024)
von: Irene Fioravanti, et al.
Veröffentlicht: (2024)
Neural Generative Models and the Parallel Architecture of Language: A Critical Review and Outlook
von: Giulia Rambelli, et al.
Veröffentlicht: (2024)
von: Giulia Rambelli, et al.
Veröffentlicht: (2024)
ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
von: Merler, Matteo, et al.
Veröffentlicht: (2025)
von: Merler, Matteo, et al.
Veröffentlicht: (2025)
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
von: Dang, Vy Tuong, et al.
Veröffentlicht: (2025)
von: Dang, Vy Tuong, et al.
Veröffentlicht: (2025)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Large language models as oracles for instantiating ontologies with domain-specific knowledge
von: Ciatto, Giovanni, et al.
Veröffentlicht: (2024)
von: Ciatto, Giovanni, et al.
Veröffentlicht: (2024)
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning
von: Yang, Bohao, et al.
Veröffentlicht: (2025)
von: Yang, Bohao, et al.
Veröffentlicht: (2025)
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary Adaptation
von: Moroni, Luca, et al.
Veröffentlicht: (2025)
von: Moroni, Luca, et al.
Veröffentlicht: (2025)
MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
von: Lazzaroni, Ruggero Marino, et al.
Veröffentlicht: (2025)
von: Lazzaroni, Ruggero Marino, et al.
Veröffentlicht: (2025)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
Domain Embeddings for Generating Complex Descriptions of Concepts in Italian Language
von: Maisto, Alessandro
Veröffentlicht: (2024)
von: Maisto, Alessandro
Veröffentlicht: (2024)
Teaching Small Language Models to Learn Logic through Meta-Learning
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
von: Capone, Luca, et al.
Veröffentlicht: (2025) -
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
von: Miliani, Martina, et al.
Veröffentlicht: (2025) -
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
von: Auriemma, Serena, et al.
Veröffentlicht: (2024) -
An Experimental Comparison of the Most Popular Approaches to Fake News Detection
von: Dell'Oglio, Pietro, et al.
Veröffentlicht: (2026) -
Enhancing Debunking Effectiveness through LLM-based Personality Adaptation
von: Dell'Oglio, Pietro, et al.
Veröffentlicht: (2026)