Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2405.11212 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916251318616064 |
|---|---|
| author | Creanga, Claudiu Dinu, Liviu Petrisor |
| author_facet | Creanga, Claudiu Dinu, Liviu Petrisor |
| contents | We used Data Maps to model and characterize the AuTexTification dataset. This provides insights about the behaviour of individual samples during training across epochs (training dynamics). We characterized the samples across 3 dimensions: confidence, variability and correctness. This shows the presence of 3 regions: easy-to-learn, ambiguous and hard-to-learn examples. We used a classic CNN architecture and found out that training the model only on a subset of ambiguous examples improves the model's out-of-distribution generalization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_11212 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Automated Text Identification Using CNN and Training Dynamics Creanga, Claudiu Dinu, Liviu Petrisor Computation and Language We used Data Maps to model and characterize the AuTexTification dataset. This provides insights about the behaviour of individual samples during training across epochs (training dynamics). We characterized the samples across 3 dimensions: confidence, variability and correctness. This shows the presence of 3 regions: easy-to-learn, ambiguous and hard-to-learn examples. We used a classic CNN architecture and found out that training the model only on a subset of ambiguous examples improves the model's out-of-distribution generalization. |
| title | Automated Text Identification Using CNN and Training Dynamics |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2405.11212 |