A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Antoine, Elie, Béchet, Frédéric, Damnati, Géraldine, Langlais, Philippe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
por: Antoine, Elie, et al.
Publicado: (2024)
por: Antoine, Elie, et al.
Publicado: (2024)
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
por: Piedboeuf, Frédéric, et al.
Publicado: (2024)
por: Piedboeuf, Frédéric, et al.
Publicado: (2024)
Increasing faithfulness in human-human dialog summarization with Spoken Language Understanding tasks
por: Akani, Eunice, et al.
Publicado: (2024)
por: Akani, Eunice, et al.
Publicado: (2024)
$\textit{BenchIE}^{FL}$ : A Manually Re-Annotated Fact-Based Open Information Extraction Benchmark
por: Lamarche, Fabrice, et al.
Publicado: (2024)
por: Lamarche, Fabrice, et al.
Publicado: (2024)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
por: Barboule, Camille, et al.
Publicado: (2024)
por: Barboule, Camille, et al.
Publicado: (2024)
Is linguistically-motivated data augmentation worth it?
por: Groshan, Ray, et al.
Publicado: (2025)
por: Groshan, Ray, et al.
Publicado: (2025)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
por: Kobeissi, Amine, et al.
Publicado: (2026)
por: Kobeissi, Amine, et al.
Publicado: (2026)
Visualizing attention zones in machine reading comprehension models
por: Cui, Yiming, et al.
Publicado: (2024)
por: Cui, Yiming, et al.
Publicado: (2024)
O_FT@EvalLLM2025 : étude comparative de choix de données et de stratégies d'apprentissage pour l'adaptation de modèles de langue à un domaine
por: Rousseau, Ismaël, et al.
Publicado: (2025)
por: Rousseau, Ismaël, et al.
Publicado: (2025)
DivMerge: A divergence-based model merging method for multi-tasking
por: Touayouch, Brahim, et al.
Publicado: (2025)
por: Touayouch, Brahim, et al.
Publicado: (2025)
WikiFactDiff: A Large, Realistic, and Temporally Adaptable Dataset for Atomic Factual Knowledge Update in Causal Language Models
por: Khodja, Hichem Ammar, et al.
Publicado: (2024)
por: Khodja, Hichem Ammar, et al.
Publicado: (2024)
On the importance of Data Scale in Pretraining Arabic Language Models
por: Ghaddar, Abbas, et al.
Publicado: (2024)
por: Ghaddar, Abbas, et al.
Publicado: (2024)
Withdrawn: Exploring the associations among task complexity, task motivation, task engagement, and linguistic complexity in L2 writing
Publicado: (2024)
Publicado: (2024)
EUROPA: A Legal Multilingual Keyphrase Generation Dataset
por: Salaün, Olivier, et al.
Publicado: (2024)
por: Salaün, Olivier, et al.
Publicado: (2024)
Statistical Deficiency for Task Inclusion Estimation
por: Fosse, Loïc, et al.
Publicado: (2025)
por: Fosse, Loïc, et al.
Publicado: (2025)
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
por: Bonzi, Doria, et al.
Publicado: (2025)
por: Bonzi, Doria, et al.
Publicado: (2025)
ReGLA: Refining Gated Linear Attention
por: Lu, Peng, et al.
Publicado: (2025)
por: Lu, Peng, et al.
Publicado: (2025)
Annotating Scientific Uncertainty: A comprehensive model using linguistic patterns and comparison with existing approaches
por: Ningrum, Panggih Kusuma, et al.
Publicado: (2025)
por: Ningrum, Panggih Kusuma, et al.
Publicado: (2025)
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
por: Khodja, Hichem Ammar, et al.
Publicado: (2025)
por: Khodja, Hichem Ammar, et al.
Publicado: (2025)
SLPL SHROOM at SemEval2024 Task 06: A comprehensive study on models ability to detect hallucination
por: Fallah, Pouya, et al.
Publicado: (2024)
por: Fallah, Pouya, et al.
Publicado: (2024)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
por: Arora, Aryaman, et al.
Publicado: (2024)
por: Arora, Aryaman, et al.
Publicado: (2024)
Power in Numbers: Robust reading comprehension by finetuning with four adversarial sentences per example
por: Marcus, Ariel
Publicado: (2024)
por: Marcus, Ariel
Publicado: (2024)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
por: Lu, Peng, et al.
Publicado: (2023)
por: Lu, Peng, et al.
Publicado: (2023)
A conclusive remark on linguistic theorizing and language modeling
por: Chesi, Cristiano
Publicado: (2025)
por: Chesi, Cristiano
Publicado: (2025)
WikiNER-fr-gold: A Gold-Standard NER Corpus
por: Cao, Danrun, et al.
Publicado: (2024)
por: Cao, Danrun, et al.
Publicado: (2024)
Leveraging language models for summarizing mental state examinations: A comprehensive evaluation and dataset release
por: Sahu, Nilesh Kumar, et al.
Publicado: (2024)
por: Sahu, Nilesh Kumar, et al.
Publicado: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
por: Ghaddar, Abbas, et al.
Publicado: (2024)
por: Ghaddar, Abbas, et al.
Publicado: (2024)
Demonstration-based learning for few-shot biomedical named entity recognition under machine reading comprehension
por: Su, Leilei, et al.
Publicado: (2023)
por: Su, Leilei, et al.
Publicado: (2023)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
por: Vázquez, Alain, et al.
Publicado: (2026)
por: Vázquez, Alain, et al.
Publicado: (2026)
Large language models and linguistic intentionality
por: Grindrod, Jumbly
Publicado: (2024)
por: Grindrod, Jumbly
Publicado: (2024)
Do language models accommodate their users? A study of linguistic convergence
por: Blevins, Terra, et al.
Publicado: (2025)
por: Blevins, Terra, et al.
Publicado: (2025)
Testing AI on language comprehension tasks reveals insensitivity to underlying meaning
por: Dentella, Vittoria, et al.
Publicado: (2023)
por: Dentella, Vittoria, et al.
Publicado: (2023)
Cutting through the noise to motivate people: A comprehensive analysis of COVID-19 social media posts de/motivating vaccination
por: Rahman, Ashiqur, et al.
Publicado: (2024)
por: Rahman, Ashiqur, et al.
Publicado: (2024)
Evaluating Polish linguistic and cultural competency in large language models
por: Dadas, Sławomir, et al.
Publicado: (2025)
por: Dadas, Sławomir, et al.
Publicado: (2025)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
por: Swain, Sankalp Tattwadarshi, et al.
Publicado: (2025)
por: Swain, Sankalp Tattwadarshi, et al.
Publicado: (2025)
A blind spot for large language models: Supradiegetic linguistic information
por: Zimmerman, Julia Witte, et al.
Publicado: (2023)
por: Zimmerman, Julia Witte, et al.
Publicado: (2023)
Identifying the sources of ideological bias in GPT models through linguistic variation in output
por: Walker, Christina, et al.
Publicado: (2024)
por: Walker, Christina, et al.
Publicado: (2024)
Merge-based syntax is mediated by distinct neurocognitive mechanisms: A clustering analysis of comprehension abilities in 84,000 individuals with language deficits across nine languages
por: Murphy, Elliot, et al.
Publicado: (2025)
por: Murphy, Elliot, et al.
Publicado: (2025)
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
por: Chizhov, Pavel, et al.
Publicado: (2025)
por: Chizhov, Pavel, et al.
Publicado: (2025)
Toxicity of the Commons: Curating Open-Source Pre-Training Data
por: Arnett, Catherine, et al.
Publicado: (2024)
por: Arnett, Catherine, et al.
Publicado: (2024)
Ejemplares similares
-
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
por: Antoine, Elie, et al.
Publicado: (2024) -
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
por: Piedboeuf, Frédéric, et al.
Publicado: (2024) -
Increasing faithfulness in human-human dialog summarization with Spoken Language Understanding tasks
por: Akani, Eunice, et al.
Publicado: (2024) -
$\textit{BenchIE}^{FL}$ : A Manually Re-Annotated Fact-Based Open Information Extraction Benchmark
por: Lamarche, Fabrice, et al.
Publicado: (2024) -
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
por: Barboule, Camille, et al.
Publicado: (2024)