Determinants of Training Corpus Size for Clinical Text Classification
Fuente:
arXiv
Guardado en:
| Autores principales: | Chaturvedi, Jaya, Deshpande, Saniya, Ma, Chenkai, Cobb, Robert, Roberts, Angus, Stewart, Robert, Stahl, Daniel, Shamsutdinova, Diana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sample Size Calculations for Developing Clinical Prediction Models: Overview and pmsims R package
por: Shamsutdinova, Diana, et al.
Publicado: (2026)
por: Shamsutdinova, Diana, et al.
Publicado: (2026)
Adaptive Gaussian Process Search for Simulation-Based Sample Size Estimation in Clinical Prediction Models: Validation of the pmsims R Package
por: Olaniran, Oyebayo Ridwan, et al.
Publicado: (2026)
por: Olaniran, Oyebayo Ridwan, et al.
Publicado: (2026)
Training LLMs over Neurally Compressed Text
por: Lester, Brian, et al.
Publicado: (2024)
por: Lester, Brian, et al.
Publicado: (2024)
Margin Discrepancy-based Adversarial Training for Multi-Domain Text Classification
por: Wu, Yuan
Publicado: (2024)
por: Wu, Yuan
Publicado: (2024)
SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas
por: Wolff, Cornelius, et al.
Publicado: (2025)
por: Wolff, Cornelius, et al.
Publicado: (2025)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
por: Pangakis, Nicholas, et al.
Publicado: (2024)
por: Pangakis, Nicholas, et al.
Publicado: (2024)
DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
por: Thorat, Shantanu, et al.
Publicado: (2025)
por: Thorat, Shantanu, et al.
Publicado: (2025)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
por: Pecher, Branislav, et al.
Publicado: (2026)
por: Pecher, Branislav, et al.
Publicado: (2026)
Learning Semantic Structure through First-Order-Logic Translation
por: Chaturvedi, Akshay, et al.
Publicado: (2024)
por: Chaturvedi, Akshay, et al.
Publicado: (2024)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
por: Schröder, Christopher, et al.
Publicado: (2024)
por: Schröder, Christopher, et al.
Publicado: (2024)
Leveraging Parameter Efficient Training Methods for Low Resource Text Classification: A Case Study in Marathi
por: Deshmukh, Pranita, et al.
Publicado: (2024)
por: Deshmukh, Pranita, et al.
Publicado: (2024)
On the Fragility of Active Learners for Text Classification
por: Ghose, Abhishek, et al.
Publicado: (2024)
por: Ghose, Abhishek, et al.
Publicado: (2024)
Universal Cross-Lingual Text Classification
por: Savant, Riya, et al.
Publicado: (2024)
por: Savant, Riya, et al.
Publicado: (2024)
Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
por: Kuo, Hsun-Yu, et al.
Publicado: (2024)
por: Kuo, Hsun-Yu, et al.
Publicado: (2024)
Retrieve, Then Classify: Corpus-Grounded Automation of Clinical Value Set Authoring
por: Mukherjee, Sumit, et al.
Publicado: (2026)
por: Mukherjee, Sumit, et al.
Publicado: (2026)
TextAge: A Curated and Diverse Text Dataset for Age Classification
por: Cheekati, Shravan, et al.
Publicado: (2024)
por: Cheekati, Shravan, et al.
Publicado: (2024)
Scaling Law for Language Models Training Considering Batch Size
por: Shuai, Xian, et al.
Publicado: (2024)
por: Shuai, Xian, et al.
Publicado: (2024)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
por: Shivagunde, Namrata, et al.
Publicado: (2026)
por: Shivagunde, Namrata, et al.
Publicado: (2026)
Controllable Synthetic Clinical Note Generation with Privacy Guarantees
por: Baumel, Tal, et al.
Publicado: (2024)
por: Baumel, Tal, et al.
Publicado: (2024)
Fair Text Classification via Transferable Representations
por: Leteno, Thibaud, et al.
Publicado: (2025)
por: Leteno, Thibaud, et al.
Publicado: (2025)
Revisiting Hierarchical Text Classification: Inference and Metrics
por: Plaud, Roman, et al.
Publicado: (2024)
por: Plaud, Roman, et al.
Publicado: (2024)
Ensembling Finetuned Language Models for Text Classification
por: Arango, Sebastian Pineda, et al.
Publicado: (2024)
por: Arango, Sebastian Pineda, et al.
Publicado: (2024)
The Moral Foundations Weibo Corpus
por: Cao, Renjie, et al.
Publicado: (2024)
por: Cao, Renjie, et al.
Publicado: (2024)
Nebula: A discourse aware Minecraft Builder
por: Chaturvedi, Akshay, et al.
Publicado: (2024)
por: Chaturvedi, Akshay, et al.
Publicado: (2024)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
por: Kesgin, H. Toprak, et al.
Publicado: (2024)
por: Kesgin, H. Toprak, et al.
Publicado: (2024)
Stochastic Adversarial Networks for Multi-Domain Text Classification
por: Wang, Xu, et al.
Publicado: (2024)
por: Wang, Xu, et al.
Publicado: (2024)
Generative or Discriminative? Revisiting Text Classification in the Era of Transformers
por: Kasa, Siva Rajesh, et al.
Publicado: (2025)
por: Kasa, Siva Rajesh, et al.
Publicado: (2025)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
por: Zhang, Ziyang, et al.
Publicado: (2025)
por: Zhang, Ziyang, et al.
Publicado: (2025)
AutoGeTS: Knowledge-based Automated Generation of Text Synthetics for Improving Text Classification
por: Xue, Chenhao, et al.
Publicado: (2025)
por: Xue, Chenhao, et al.
Publicado: (2025)
MatchXML: An Efficient Text-label Matching Framework for Extreme Multi-label Text Classification
por: Ye, Hui, et al.
Publicado: (2023)
por: Ye, Hui, et al.
Publicado: (2023)
Bridging Electronic Health Records and Clinical Texts: Contrastive Learning for Enhanced Clinical Tasks
por: Ketabi, Sara, et al.
Publicado: (2025)
por: Ketabi, Sara, et al.
Publicado: (2025)
Cheap Ways of Extracting Clinical Markers from Texts
por: Sandu, Anastasia, et al.
Publicado: (2024)
por: Sandu, Anastasia, et al.
Publicado: (2024)
Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer Approach
por: Chaturvedi, Rochana, et al.
Publicado: (2025)
por: Chaturvedi, Rochana, et al.
Publicado: (2025)
Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning
por: Rahman, Md Mushfiqur, et al.
Publicado: (2024)
por: Rahman, Md Mushfiqur, et al.
Publicado: (2024)
How to Train Text Summarization Model with Weak Supervisions
por: Wang, Yanbo, et al.
Publicado: (2024)
por: Wang, Yanbo, et al.
Publicado: (2024)
Hate Speech Detection and Classification in Amharic Text with Deep Learning
por: Gashe, Samuel Minale, et al.
Publicado: (2024)
por: Gashe, Samuel Minale, et al.
Publicado: (2024)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
por: Eshuijs, Leon, et al.
Publicado: (2025)
por: Eshuijs, Leon, et al.
Publicado: (2025)
Evaluating Text Classification Robustness to Part-of-Speech Adversarial Examples
por: Samadi, Anahita, et al.
Publicado: (2024)
por: Samadi, Anahita, et al.
Publicado: (2024)
AstroConcepts: A Large-Scale Multi-Label Classification Corpus for Astrophysics
por: Alkan, Atilla Kaan, et al.
Publicado: (2026)
por: Alkan, Atilla Kaan, et al.
Publicado: (2026)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
por: Gupta, Aarav, et al.
Publicado: (2026)
por: Gupta, Aarav, et al.
Publicado: (2026)
Ejemplares similares
-
Sample Size Calculations for Developing Clinical Prediction Models: Overview and pmsims R package
por: Shamsutdinova, Diana, et al.
Publicado: (2026) -
Adaptive Gaussian Process Search for Simulation-Based Sample Size Estimation in Clinical Prediction Models: Validation of the pmsims R Package
por: Olaniran, Oyebayo Ridwan, et al.
Publicado: (2026) -
Training LLMs over Neurally Compressed Text
por: Lester, Brian, et al.
Publicado: (2024) -
Margin Discrepancy-based Adversarial Training for Multi-Domain Text Classification
por: Wu, Yuan
Publicado: (2024) -
SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas
por: Wolff, Cornelius, et al.
Publicado: (2025)