Investigating a Benchmark for Training-set free Evaluation of Linguistic Capabilities in Machine Reading Comprehension
Fuente:
arXiv
Guardado en:
| Autores principales: | Schlegel, Viktor, Nenadic, Goran, Batista-Navarro, Riza |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
por: Wu, Yulong, et al.
Publicado: (2025)
por: Wu, Yulong, et al.
Publicado: (2025)
Large Language Models in Argument Mining: A Survey
por: Li, Hao, et al.
Publicado: (2025)
por: Li, Hao, et al.
Publicado: (2025)
Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement
por: Li, Hao, et al.
Publicado: (2025)
por: Li, Hao, et al.
Publicado: (2025)
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
por: Hong, Kung Yin, et al.
Publicado: (2024)
por: Hong, Kung Yin, et al.
Publicado: (2024)
CantonMT: Cantonese to English NMT Platform with Fine-Tuned Models Using Synthetic Back-Translation Data
por: Hong, Kung Yin, et al.
Publicado: (2024)
por: Hong, Kung Yin, et al.
Publicado: (2024)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
por: Wu, Yulong, et al.
Publicado: (2025)
por: Wu, Yulong, et al.
Publicado: (2025)
BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion Modeling
por: Li, Hao, et al.
Publicado: (2025)
por: Li, Hao, et al.
Publicado: (2025)
Which Side Are You On? A Multi-task Dataset for End-to-End Argument Summarisation and Evaluation
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
por: Feher, Darius, et al.
Publicado: (2024)
por: Feher, Darius, et al.
Publicado: (2024)
TriG-NER: Triplet-Grid Framework for Discontinuous Named Entity Recognition
por: Cabral, Rina Carines, et al.
Publicado: (2024)
por: Cabral, Rina Carines, et al.
Publicado: (2024)
Structured Information Matters: Explainable ICD Coding with Patient-Level Knowledge Graphs
por: Li, Mingyang, et al.
Publicado: (2025)
por: Li, Mingyang, et al.
Publicado: (2025)
Aspect-based Sentiment Evaluation of Chess Moves (ASSESS): an NLP-based Method for Evaluating Chess Strategies from Textbooks
por: Alrdahi, Haifa, et al.
Publicado: (2024)
por: Alrdahi, Haifa, et al.
Publicado: (2024)
Neural Machine Translation of Clinical Text: An Empirical Investigation into Multilingual Pre-Trained Language Models and Transfer-Learning
por: Han, Lifeng, et al.
Publicado: (2023)
por: Han, Lifeng, et al.
Publicado: (2023)
AutoLLM-CARD: Towards a Description and Landscape of Large Language Models
por: Tian, Shengwei, et al.
Publicado: (2024)
por: Tian, Shengwei, et al.
Publicado: (2024)
M-QALM: A Benchmark to Assess Clinical Reading Comprehension and Knowledge Recall in Large Language Models via Question Answering
por: Subramanian, Anand, et al.
Publicado: (2024)
por: Subramanian, Anand, et al.
Publicado: (2024)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
por: Madusanka, Tharindu, et al.
Publicado: (2025)
por: Madusanka, Tharindu, et al.
Publicado: (2025)
INSIGHTBUDDY-AI: Medication Extraction and Entity Linking using Large Language Models and Ensemble Learning
por: Romero, Pablo, et al.
Publicado: (2024)
por: Romero, Pablo, et al.
Publicado: (2024)
MTUncertainty: Assessing the Need for Post-editing of Machine Translation Outputs by Fine-tuning OpenAI LLMs
por: Gladkoff, Serge, et al.
Publicado: (2023)
por: Gladkoff, Serge, et al.
Publicado: (2023)
De-identification of clinical free text using natural language processing: A systematic review of current approaches
por: Kovačević, Aleksandar, et al.
Publicado: (2023)
por: Kovačević, Aleksandar, et al.
Publicado: (2023)
Investigating Large Language Models and Control Mechanisms to Improve Text Readability of Biomedical Abstracts
por: Li, Zihao, et al.
Publicado: (2023)
por: Li, Zihao, et al.
Publicado: (2023)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
por: Ma, Shengkun, et al.
Publicado: (2025)
por: Ma, Shengkun, et al.
Publicado: (2025)
Investigating Recent Large Language Models for Vietnamese Machine Reading Comprehension
por: Nguyen, Anh Duc, et al.
Publicado: (2025)
por: Nguyen, Anh Duc, et al.
Publicado: (2025)
Extract-and-Abstract: Unifying Extractive and Abstractive Summarization within Single Encoder-Decoder Framework
por: Wu, Yuping, et al.
Publicado: (2024)
por: Wu, Yuping, et al.
Publicado: (2024)
Will Large Language Models Transform Clinical Prediction?
por: Yildiz, Yusuf, et al.
Publicado: (2025)
por: Yildiz, Yusuf, et al.
Publicado: (2025)
Exploration of Masked and Causal Language Modelling for Text Generation
por: Micheletti, Nicolo, et al.
Publicado: (2024)
por: Micheletti, Nicolo, et al.
Publicado: (2024)
A Comparative Study on Automatic Coding of Medical Letters with Explainability
por: Glen, Jamie, et al.
Publicado: (2024)
por: Glen, Jamie, et al.
Publicado: (2024)
MaLei at the PLABA Track of TREC 2024: RoBERTa for Term Replacement -- LLaMA3.1 and GPT-4o for Complete Abstract Adaptation
por: Ling, Zhidong, et al.
Publicado: (2024)
por: Ling, Zhidong, et al.
Publicado: (2024)
HealthcareNLP: where are we and what is next?
por: Han, Lifeng, et al.
Publicado: (2025)
por: Han, Lifeng, et al.
Publicado: (2025)
Evaluating the Robustness of Machine Reading Comprehension Models to Low Resource Entity Renaming
por: Siro, Clemencia, et al.
Publicado: (2023)
por: Siro, Clemencia, et al.
Publicado: (2023)
JobResQA: A Benchmark for LLM Machine Reading Comprehension on Multilingual Résumés and JDs
por: Carrino, Casimiro Pio, et al.
Publicado: (2026)
por: Carrino, Casimiro Pio, et al.
Publicado: (2026)
Generating Synthetic Free-text Medical Records with Low Re-identification Risk using Masked Language Modeling
por: Belkadi, Samuel, et al.
Publicado: (2024)
por: Belkadi, Samuel, et al.
Publicado: (2024)
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
por: Wu, Siwei, et al.
Publicado: (2025)
por: Wu, Siwei, et al.
Publicado: (2025)
Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading Comprehension
por: Ai, Lin, et al.
Publicado: (2024)
por: Ai, Lin, et al.
Publicado: (2024)
Synthetic4Health: Generating Annotated Synthetic Clinical Letters
por: Ren, Libo, et al.
Publicado: (2024)
por: Ren, Libo, et al.
Publicado: (2024)
Automated Boilerplate: Prevalence and Quality of Contract Generators in the Context of Swiss Privacy Policies
por: Nenadic, Luka, et al.
Publicado: (2025)
por: Nenadic, Luka, et al.
Publicado: (2025)
DeIDClinic: A Risk-Aware Pseudonymization Framework for Clinical Text De-identification and Re-identification Risk Assessment
por: Paul, Angel, et al.
Publicado: (2024)
por: Paul, Angel, et al.
Publicado: (2024)
Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
por: AlMannaa, Feras, et al.
Publicado: (2025)
por: AlMannaa, Feras, et al.
Publicado: (2025)
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
por: Sun, Yizheng, et al.
Publicado: (2025)
por: Sun, Yizheng, et al.
Publicado: (2025)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
por: Dogan, Mustafa, et al.
Publicado: (2024)
por: Dogan, Mustafa, et al.
Publicado: (2024)
Large Language Models for Biomedical Text Simplification: Promising But Not There Yet
por: Li, Zihao, et al.
Publicado: (2024)
por: Li, Zihao, et al.
Publicado: (2024)
Ejemplares similares
-
Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
por: Wu, Yulong, et al.
Publicado: (2025) -
Large Language Models in Argument Mining: A Survey
por: Li, Hao, et al.
Publicado: (2025) -
Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement
por: Li, Hao, et al.
Publicado: (2025) -
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
por: Hong, Kung Yin, et al.
Publicado: (2024) -
CantonMT: Cantonese to English NMT Platform with Fine-Tuned Models Using Synthetic Back-Translation Data
por: Hong, Kung Yin, et al.
Publicado: (2024)