Determinants of Training Corpus Size for Clinical Text Classification
Fuente:
arXiv
Salvato in:
| Autori principali: | Chaturvedi, Jaya, Deshpande, Saniya, Ma, Chenkai, Cobb, Robert, Roberts, Angus, Stewart, Robert, Stahl, Daniel, Shamsutdinova, Diana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sample Size Calculations for Developing Clinical Prediction Models: Overview and pmsims R package
di: Shamsutdinova, Diana, et al.
Pubblicazione: (2026)
di: Shamsutdinova, Diana, et al.
Pubblicazione: (2026)
Adaptive Gaussian Process Search for Simulation-Based Sample Size Estimation in Clinical Prediction Models: Validation of the pmsims R Package
di: Olaniran, Oyebayo Ridwan, et al.
Pubblicazione: (2026)
di: Olaniran, Oyebayo Ridwan, et al.
Pubblicazione: (2026)
Training LLMs over Neurally Compressed Text
di: Lester, Brian, et al.
Pubblicazione: (2024)
di: Lester, Brian, et al.
Pubblicazione: (2024)
Margin Discrepancy-based Adversarial Training for Multi-Domain Text Classification
di: Wu, Yuan
Pubblicazione: (2024)
di: Wu, Yuan
Pubblicazione: (2024)
SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas
di: Wolff, Cornelius, et al.
Pubblicazione: (2025)
di: Wolff, Cornelius, et al.
Pubblicazione: (2025)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
di: Thorat, Shantanu, et al.
Pubblicazione: (2025)
di: Thorat, Shantanu, et al.
Pubblicazione: (2025)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
di: Pecher, Branislav, et al.
Pubblicazione: (2026)
di: Pecher, Branislav, et al.
Pubblicazione: (2026)
Learning Semantic Structure through First-Order-Logic Translation
di: Chaturvedi, Akshay, et al.
Pubblicazione: (2024)
di: Chaturvedi, Akshay, et al.
Pubblicazione: (2024)
Leveraging Parameter Efficient Training Methods for Low Resource Text Classification: A Case Study in Marathi
di: Deshmukh, Pranita, et al.
Pubblicazione: (2024)
di: Deshmukh, Pranita, et al.
Pubblicazione: (2024)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
di: Schröder, Christopher, et al.
Pubblicazione: (2024)
di: Schröder, Christopher, et al.
Pubblicazione: (2024)
On the Fragility of Active Learners for Text Classification
di: Ghose, Abhishek, et al.
Pubblicazione: (2024)
di: Ghose, Abhishek, et al.
Pubblicazione: (2024)
Universal Cross-Lingual Text Classification
di: Savant, Riya, et al.
Pubblicazione: (2024)
di: Savant, Riya, et al.
Pubblicazione: (2024)
Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
di: Kuo, Hsun-Yu, et al.
Pubblicazione: (2024)
di: Kuo, Hsun-Yu, et al.
Pubblicazione: (2024)
TextAge: A Curated and Diverse Text Dataset for Age Classification
di: Cheekati, Shravan, et al.
Pubblicazione: (2024)
di: Cheekati, Shravan, et al.
Pubblicazione: (2024)
Retrieve, Then Classify: Corpus-Grounded Automation of Clinical Value Set Authoring
di: Mukherjee, Sumit, et al.
Pubblicazione: (2026)
di: Mukherjee, Sumit, et al.
Pubblicazione: (2026)
Scaling Law for Language Models Training Considering Batch Size
di: Shuai, Xian, et al.
Pubblicazione: (2024)
di: Shuai, Xian, et al.
Pubblicazione: (2024)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
Controllable Synthetic Clinical Note Generation with Privacy Guarantees
di: Baumel, Tal, et al.
Pubblicazione: (2024)
di: Baumel, Tal, et al.
Pubblicazione: (2024)
Fair Text Classification via Transferable Representations
di: Leteno, Thibaud, et al.
Pubblicazione: (2025)
di: Leteno, Thibaud, et al.
Pubblicazione: (2025)
Revisiting Hierarchical Text Classification: Inference and Metrics
di: Plaud, Roman, et al.
Pubblicazione: (2024)
di: Plaud, Roman, et al.
Pubblicazione: (2024)
Ensembling Finetuned Language Models for Text Classification
di: Arango, Sebastian Pineda, et al.
Pubblicazione: (2024)
di: Arango, Sebastian Pineda, et al.
Pubblicazione: (2024)
The Moral Foundations Weibo Corpus
di: Cao, Renjie, et al.
Pubblicazione: (2024)
di: Cao, Renjie, et al.
Pubblicazione: (2024)
Nebula: A discourse aware Minecraft Builder
di: Chaturvedi, Akshay, et al.
Pubblicazione: (2024)
di: Chaturvedi, Akshay, et al.
Pubblicazione: (2024)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
Stochastic Adversarial Networks for Multi-Domain Text Classification
di: Wang, Xu, et al.
Pubblicazione: (2024)
di: Wang, Xu, et al.
Pubblicazione: (2024)
Generative or Discriminative? Revisiting Text Classification in the Era of Transformers
di: Kasa, Siva Rajesh, et al.
Pubblicazione: (2025)
di: Kasa, Siva Rajesh, et al.
Pubblicazione: (2025)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
AutoGeTS: Knowledge-based Automated Generation of Text Synthetics for Improving Text Classification
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
MatchXML: An Efficient Text-label Matching Framework for Extreme Multi-label Text Classification
di: Ye, Hui, et al.
Pubblicazione: (2023)
di: Ye, Hui, et al.
Pubblicazione: (2023)
Bridging Electronic Health Records and Clinical Texts: Contrastive Learning for Enhanced Clinical Tasks
di: Ketabi, Sara, et al.
Pubblicazione: (2025)
di: Ketabi, Sara, et al.
Pubblicazione: (2025)
Cheap Ways of Extracting Clinical Markers from Texts
di: Sandu, Anastasia, et al.
Pubblicazione: (2024)
di: Sandu, Anastasia, et al.
Pubblicazione: (2024)
Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer Approach
di: Chaturvedi, Rochana, et al.
Pubblicazione: (2025)
di: Chaturvedi, Rochana, et al.
Pubblicazione: (2025)
Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning
di: Rahman, Md Mushfiqur, et al.
Pubblicazione: (2024)
di: Rahman, Md Mushfiqur, et al.
Pubblicazione: (2024)
How to Train Text Summarization Model with Weak Supervisions
di: Wang, Yanbo, et al.
Pubblicazione: (2024)
di: Wang, Yanbo, et al.
Pubblicazione: (2024)
Hate Speech Detection and Classification in Amharic Text with Deep Learning
di: Gashe, Samuel Minale, et al.
Pubblicazione: (2024)
di: Gashe, Samuel Minale, et al.
Pubblicazione: (2024)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
Evaluating Text Classification Robustness to Part-of-Speech Adversarial Examples
di: Samadi, Anahita, et al.
Pubblicazione: (2024)
di: Samadi, Anahita, et al.
Pubblicazione: (2024)
AstroConcepts: A Large-Scale Multi-Label Classification Corpus for Astrophysics
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2026)
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2026)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
di: Gupta, Aarav, et al.
Pubblicazione: (2026)
di: Gupta, Aarav, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Sample Size Calculations for Developing Clinical Prediction Models: Overview and pmsims R package
di: Shamsutdinova, Diana, et al.
Pubblicazione: (2026) -
Adaptive Gaussian Process Search for Simulation-Based Sample Size Estimation in Clinical Prediction Models: Validation of the pmsims R Package
di: Olaniran, Oyebayo Ridwan, et al.
Pubblicazione: (2026) -
Training LLMs over Neurally Compressed Text
di: Lester, Brian, et al.
Pubblicazione: (2024) -
Margin Discrepancy-based Adversarial Training for Multi-Domain Text Classification
di: Wu, Yuan
Pubblicazione: (2024) -
SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas
di: Wolff, Cornelius, et al.
Pubblicazione: (2025)