Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917030305726464 |
|---|---|
| author | Liu, Jiaqi Wang, Tong Liu, Su Hu, Xin Tong, Ran Wang, Lanruo Xu, Jiexi |
| author_facet | Liu, Jiaqi Wang, Tong Liu, Su Hu, Xin Tong, Ran Wang, Lanruo Xu, Jiexi |
| contents | The research evaluates lightweight medical abstract classification methods to establish their maximum performance capabilities under financial budget restrictions. On the public medical abstracts corpus, we finetune BERT base and Distil BERT with three objectives cross entropy (CE), class weighted CE, and focal loss under identical tokenization, sequence length, optimizer, and schedule. DistilBERT with plain CE gives the strongest raw argmax trade off, while a post hoc operating point selection (validation calibrated, classwise thresholds) sub stantially improves deployed performance; under this tuned regime, focal benefits most. We report Accuracy, Macro F1, and WeightedF1, release evaluation artifacts, and include confusion analyses to clarify error structure. The practical takeaway is to start with a compact encoder and CE, then add lightweight calibration or thresholding when deployment requires higher macro balance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_10025 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default Liu, Jiaqi Wang, Tong Liu, Su Hu, Xin Tong, Ran Wang, Lanruo Xu, Jiexi Computation and Language Artificial Intelligence The research evaluates lightweight medical abstract classification methods to establish their maximum performance capabilities under financial budget restrictions. On the public medical abstracts corpus, we finetune BERT base and Distil BERT with three objectives cross entropy (CE), class weighted CE, and focal loss under identical tokenization, sequence length, optimizer, and schedule. DistilBERT with plain CE gives the strongest raw argmax trade off, while a post hoc operating point selection (validation calibrated, classwise thresholds) sub stantially improves deployed performance; under this tuned regime, focal benefits most. We report Accuracy, Macro F1, and WeightedF1, release evaluation artifacts, and include confusion analyses to clarify error structure. The practical takeaway is to start with a compact encoder and CE, then add lightweight calibration or thresholding when deployment requires higher macro balance. |
| title | Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.10025 |