Diagnosing our datasets: How does my language model learn clinical information?
Fuente:
arXiv
Guardado en:
| Autores principales: | Jia, Furong, Sontag, David, Agrawal, Monica |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What Patients Really Ask: Exploring the Effect of False Assumptions in Patient Information Seeking
por: Xiong, Raymond, et al.
Publicado: (2026)
por: Xiong, Raymond, et al.
Publicado: (2026)
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
por: Jia, Furong, et al.
Publicado: (2025)
por: Jia, Furong, et al.
Publicado: (2025)
How does fine-tuning improve sensorimotor representations in large language models?
por: Wu, Minghua, et al.
Publicado: (2026)
por: Wu, Minghua, et al.
Publicado: (2026)
LLM Dataset Inference: Did you train on my dataset?
por: Maini, Pratyush, et al.
Publicado: (2024)
por: Maini, Pratyush, et al.
Publicado: (2024)
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models
por: Hegselmann, Stefan, et al.
Publicado: (2024)
por: Hegselmann, Stefan, et al.
Publicado: (2024)
What does it mean to understand language?
por: Casto, Colton, et al.
Publicado: (2025)
por: Casto, Colton, et al.
Publicado: (2025)
Large language models have learned to use language
por: Lupyan, Gary
Publicado: (2025)
por: Lupyan, Gary
Publicado: (2025)
How do language models learn facts? Dynamics, curricula and hallucinations
por: Zucchet, Nicolas, et al.
Publicado: (2025)
por: Zucchet, Nicolas, et al.
Publicado: (2025)
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
por: Das, Dipto, et al.
Publicado: (2025)
por: Das, Dipto, et al.
Publicado: (2025)
Context informs pragmatic interpretation in vision-language models
por: Tan, Alvin Wei Ming, et al.
Publicado: (2025)
por: Tan, Alvin Wei Ming, et al.
Publicado: (2025)
"What's my model inside of?": Exploring the role of environments for grounded natural language understanding
por: Tamari, Ronen
Publicado: (2024)
por: Tamari, Ronen
Publicado: (2024)
How much do language models memorize?
por: Morris, John X., et al.
Publicado: (2025)
por: Morris, John X., et al.
Publicado: (2025)
Infusing clinical knowledge into tokenisers for language models
por: Hasan, Abul, et al.
Publicado: (2024)
por: Hasan, Abul, et al.
Publicado: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
por: Li, Kunning, et al.
Publicado: (2025)
por: Li, Kunning, et al.
Publicado: (2025)
Dynamic data sampler for cross-language transfer learning in large language models
por: Li, Yudong, et al.
Publicado: (2024)
por: Li, Yudong, et al.
Publicado: (2024)
Efficient extraction of medication information from clinical notes: an evaluation in two languages
por: Fabacher, Thibaut, et al.
Publicado: (2025)
por: Fabacher, Thibaut, et al.
Publicado: (2025)
Theoretical Analysis of Weak-to-Strong Generalization
por: Lang, Hunter, et al.
Publicado: (2024)
por: Lang, Hunter, et al.
Publicado: (2024)
MedMobile: A mobile-sized language model with clinical capabilities
por: Vishwanath, Krithik, et al.
Publicado: (2024)
por: Vishwanath, Krithik, et al.
Publicado: (2024)
Symbol tuning improves in-context learning in language models
por: Wei, Jerry, et al.
Publicado: (2023)
por: Wei, Jerry, et al.
Publicado: (2023)
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
por: Ichmoukhamedov, Timour, et al.
Publicado: (2024)
por: Ichmoukhamedov, Timour, et al.
Publicado: (2024)
How does Misinformation Affect Large Language Model Behaviors and Preferences?
por: Peng, Miao, et al.
Publicado: (2025)
por: Peng, Miao, et al.
Publicado: (2025)
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation
por: Carandang, Kristine Ann M., et al.
Publicado: (2025)
por: Carandang, Kristine Ann M., et al.
Publicado: (2025)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
por: He, Jie, et al.
Publicado: (2025)
por: He, Jie, et al.
Publicado: (2025)
Leveraging language models for summarizing mental state examinations: A comprehensive evaluation and dataset release
por: Sahu, Nilesh Kumar, et al.
Publicado: (2024)
por: Sahu, Nilesh Kumar, et al.
Publicado: (2024)
An information-theoretic model of shallow and deep language comprehension
por: Li, Jiaxuan, et al.
Publicado: (2024)
por: Li, Jiaxuan, et al.
Publicado: (2024)
Building English ASR model with regional language support
por: Agrawal, Purvi, et al.
Publicado: (2025)
por: Agrawal, Purvi, et al.
Publicado: (2025)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
por: Matos, João, et al.
Publicado: (2024)
por: Matos, João, et al.
Publicado: (2024)
Streamlining evidence based clinical recommendations with large language models
por: Li, Dubai, et al.
Publicado: (2025)
por: Li, Dubai, et al.
Publicado: (2025)
Automatic detection of diseases in Spanish clinical notes combining medical language models and ontologies
por: Torre, Leon-Paul Schaub, et al.
Publicado: (2024)
por: Torre, Leon-Paul Schaub, et al.
Publicado: (2024)
From RAGs to riches: Utilizing large language models to write documents for clinical trials
por: Markey, Nigel, et al.
Publicado: (2024)
por: Markey, Nigel, et al.
Publicado: (2024)
The use of large language models to enhance cancer clinical trial educational materials
por: Gao, Mingye, et al.
Publicado: (2024)
por: Gao, Mingye, et al.
Publicado: (2024)
What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models
por: Verhoef, Tessa, et al.
Publicado: (2024)
por: Verhoef, Tessa, et al.
Publicado: (2024)
Toxic language detection: a systematic review of Arabic datasets
por: Bensalem, Imene, et al.
Publicado: (2023)
por: Bensalem, Imene, et al.
Publicado: (2023)
Uncovering inequalities in new knowledge learning by large language models across different languages
por: Wang, Chenglong, et al.
Publicado: (2025)
por: Wang, Chenglong, et al.
Publicado: (2025)
Do large language models resemble humans in language use?
por: Cai, Zhenguang G., et al.
Publicado: (2023)
por: Cai, Zhenguang G., et al.
Publicado: (2023)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
por: Feng, Jiarui, et al.
Publicado: (2024)
por: Feng, Jiarui, et al.
Publicado: (2024)
Meta predictive learning model of languages in neural circuits
por: Li, Chan, et al.
Publicado: (2023)
por: Li, Chan, et al.
Publicado: (2023)
RigoChat 2: an adapted language model to Spanish using a bounded dataset and reduced hardware
por: Gómez, Gonzalo Santamaría, et al.
Publicado: (2025)
por: Gómez, Gonzalo Santamaría, et al.
Publicado: (2025)
How does a Language-Specific Tokenizer affect LLMs?
por: Seo, Jean, et al.
Publicado: (2025)
por: Seo, Jean, et al.
Publicado: (2025)
Few-shot clinical entity recognition in English, French and Spanish: masked language models outperform generative model prompting
por: Naguib, Marco, et al.
Publicado: (2024)
por: Naguib, Marco, et al.
Publicado: (2024)
Ejemplares similares
-
What Patients Really Ask: Exploring the Effect of False Assumptions in Patient Information Seeking
por: Xiong, Raymond, et al.
Publicado: (2026) -
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
por: Jia, Furong, et al.
Publicado: (2025) -
How does fine-tuning improve sensorimotor representations in large language models?
por: Wu, Minghua, et al.
Publicado: (2026) -
LLM Dataset Inference: Did you train on my dataset?
por: Maini, Pratyush, et al.
Publicado: (2024) -
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models
por: Hegselmann, Stefan, et al.
Publicado: (2024)