Low-resource Information Extraction with the European Clinical Case Corpus
Fuente:
arXiv
Saved in:
| Main Authors: | Ghosh, Soumitra, Altuna, Begona, Farzi, Saeed, Ferrazzi, Pietro, Lavelli, Alberto, Mezzanotte, Giulia, Speranza, Manuela, Magnini, Bernardo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Converting Annotated Clinical Cases into Structured Case Report Forms
by: Ferrazzi, Pietro, et al.
Published: (2025)
by: Ferrazzi, Pietro, et al.
Published: (2025)
Small LLMs for Medical NLP: a Systematic Analysis of Few-Shot, Constraint Decoding, Fine-Tuning and Continual Pre-Training in Italian
by: Ferrazzi, Pietro, et al.
Published: (2026)
by: Ferrazzi, Pietro, et al.
Published: (2026)
Toward Automatic Filling of Case Report Forms: A Case Study on Data from an Italian Emergency Department
by: Kaczmarek, Gabriela Anna, et al.
Published: (2026)
by: Kaczmarek, Gabriela Anna, et al.
Published: (2026)
Evaluating Task-Oriented Dialogue Consistency through Constraint Satisfaction
by: Labruna, Tiziano, et al.
Published: (2024)
by: Labruna, Tiziano, et al.
Published: (2024)
EusHeidelTime: Time Expression Extraction and Normalisation for Basque
by: Begoña Altuna
Published: (2017)
by: Begoña Altuna
Published: (2017)
NoticIA: A Clickbait Article Summarization Dataset in Spanish
by: García-Ferrero, Iker, et al.
Published: (2024)
by: García-Ferrero, Iker, et al.
Published: (2024)
Multilingual Medical Reasoning for Question Answering with Large Language Models
by: Ferrazzi, Pietro, et al.
Published: (2025)
by: Ferrazzi, Pietro, et al.
Published: (2025)
Is Agentic RAG worth it? An experimental comparison of RAG approaches
by: Ferrazzi, Pietro, et al.
Published: (2026)
by: Ferrazzi, Pietro, et al.
Published: (2026)
EthioMT: Parallel Corpus for Low-resource Ethiopian Languages
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
by: Barnes, Jeremy, et al.
Published: (2025)
by: Barnes, Jeremy, et al.
Published: (2025)
CHisIEC: An Information Extraction Corpus for Ancient Chinese History
by: Tang, Xuemei, et al.
Published: (2024)
by: Tang, Xuemei, et al.
Published: (2024)
A Hard Nut to Crack: Idiom Detection with Conversational Large Language Models
by: Fornaciari, Francesca De Luca, et al.
Published: (2024)
by: Fornaciari, Francesca De Luca, et al.
Published: (2024)
ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
by: Merler, Matteo, et al.
Published: (2025)
by: Merler, Matteo, et al.
Published: (2025)
Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
by: García-Ferrero, Iker, et al.
Published: (2024)
by: García-Ferrero, Iker, et al.
Published: (2024)
Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows
by: Sebe, Clémence, et al.
Published: (2024)
by: Sebe, Clémence, et al.
Published: (2024)
Evalita-LLM: Benchmarking Large Language Models on Italian
by: Magnini, Bernardo, et al.
Published: (2025)
by: Magnini, Bernardo, et al.
Published: (2025)
SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts
by: Lasandi, Minduli, et al.
Published: (2026)
by: Lasandi, Minduli, et al.
Published: (2026)
LASQ: A Low-resource Aspect-based Sentiment Quadruple Extraction Dataset
by: Yusufu, Aizihaierjiang, et al.
Published: (2026)
by: Yusufu, Aizihaierjiang, et al.
Published: (2026)
CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports
by: Zhang, Xiao Yu Cindy, et al.
Published: (2025)
by: Zhang, Xiao Yu Cindy, et al.
Published: (2025)
Supporting Humans in Evaluating AI Summaries of Legal Depositions
by: Farzi, Naghmeh, et al.
Published: (2026)
by: Farzi, Naghmeh, et al.
Published: (2026)
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus
by: Singh, Pooja, et al.
Published: (2025)
by: Singh, Pooja, et al.
Published: (2025)
Event-Arguments Extraction Corpus and Modeling using BERT for Arabic
by: Aljabari, Alaa, et al.
Published: (2024)
by: Aljabari, Alaa, et al.
Published: (2024)
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark
by: Testa, Davide, et al.
Published: (2025)
by: Testa, Davide, et al.
Published: (2025)
IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
by: Paul, Shounak, et al.
Published: (2025)
by: Paul, Shounak, et al.
Published: (2025)
ACE-2005-PT: Corpus for Event Extraction in Portuguese
by: Cunha, Luís Filipe, et al.
Published: (2024)
by: Cunha, Luís Filipe, et al.
Published: (2024)
IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus
by: Gui, Honghao, et al.
Published: (2024)
by: Gui, Honghao, et al.
Published: (2024)
Just a Scratch: Enhancing LLM Capabilities for Self-harm Detection through Intent Differentiation and Emoji Interpretation
by: Ghosh, Soumitra, et al.
Published: (2025)
by: Ghosh, Soumitra, et al.
Published: (2025)
Small Language Models for Privacy-Preserving Clinical Information Extraction in Low-Resource Languages
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2026)
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2026)
A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language Models
by: Park, Namu, et al.
Published: (2024)
by: Park, Namu, et al.
Published: (2024)
Introducing Syllable Tokenization for Low-resource Languages: A Case Study with Swahili
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
CoastTerm: a Corpus for Multidisciplinary Term Extraction in Coastal Scientific Literature
by: Delaunay, Julien, et al.
Published: (2024)
by: Delaunay, Julien, et al.
Published: (2024)
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
by: Teixeira, Francisco, et al.
Published: (2026)
by: Teixeira, Francisco, et al.
Published: (2026)
AXE: Low-Cost Cross-Domain Web Structured Information Extraction
by: Mansour, Abdelrahman, et al.
Published: (2026)
by: Mansour, Abdelrahman, et al.
Published: (2026)
Determinants of Training Corpus Size for Clinical Text Classification
by: Chaturvedi, Jaya, et al.
Published: (2026)
by: Chaturvedi, Jaya, et al.
Published: (2026)
An Efficient Approach for Machine Translation on Low-resource Languages: A Case Study in Vietnamese-Chinese
by: Son, Tran Ngoc, et al.
Published: (2025)
by: Son, Tran Ngoc, et al.
Published: (2025)
STARK: Spatio-Temporal Attention for Representation of Keypoints for Continuous Sign Language Recognition
by: Patra, Suvajit, et al.
Published: (2026)
by: Patra, Suvajit, et al.
Published: (2026)
Large language models as oracles for instantiating ontologies with domain-specific knowledge
by: Ciatto, Giovanni, et al.
Published: (2024)
by: Ciatto, Giovanni, et al.
Published: (2024)
PumpSense: Real-Time Detection and Target Extraction of Crypto Pump-and-Dumps on Telegram
by: Mahrous, Ahmed, et al.
Published: (2026)
by: Mahrous, Ahmed, et al.
Published: (2026)
PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text
by: Kunle-John, Ifeoluwa, et al.
Published: (2026)
by: Kunle-John, Ifeoluwa, et al.
Published: (2026)
Similar Items
-
Converting Annotated Clinical Cases into Structured Case Report Forms
by: Ferrazzi, Pietro, et al.
Published: (2025) -
Small LLMs for Medical NLP: a Systematic Analysis of Few-Shot, Constraint Decoding, Fine-Tuning and Continual Pre-Training in Italian
by: Ferrazzi, Pietro, et al.
Published: (2026) -
Toward Automatic Filling of Case Report Forms: A Case Study on Data from an Italian Emergency Department
by: Kaczmarek, Gabriela Anna, et al.
Published: (2026) -
Evaluating Task-Oriented Dialogue Consistency through Constraint Satisfaction
by: Labruna, Tiziano, et al.
Published: (2024) -
EusHeidelTime: Time Expression Extraction and Normalisation for Basque
by: Begoña Altuna
Published: (2017)