Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.01006 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912357638209536 |
|---|---|
| author | Mamtani, Sumit Sonawane, Maitreya Agarwal, Kanika Sanjeev, Nishanth |
| author_facet | Mamtani, Sumit Sonawane, Maitreya Agarwal, Kanika Sanjeev, Nishanth |
| contents | Tokenization is a foundational step in most natural language processing (NLP) pipelines, yet it introduces challenges such as vocabulary mismatch and out-of-vocabulary issues. Recent work has shown that models operating directly on raw text at the byte or character level can mitigate these limitations. In this paper, we evaluate two token-free models, ByT5 and CANINE, on the task of sarcasm detection in both social media (Twitter) and non-social media (news headlines) domains. We fine-tune and benchmark these models against token-based baselines and state-of-the-art approaches. Our results show that ByT5-small and CANINE outperform token-based counterparts and achieve new state-of-the-art performance, improving accuracy by 0.77% and 0.49% on the News Headlines and Twitter Sarcasm datasets, respectively. These findings underscore the potential of token-free models for robust NLP in noisy and informal domains such as social media. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_01006 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Token-free Models for Sarcasm Detection Mamtani, Sumit Sonawane, Maitreya Agarwal, Kanika Sanjeev, Nishanth Computation and Language Tokenization is a foundational step in most natural language processing (NLP) pipelines, yet it introduces challenges such as vocabulary mismatch and out-of-vocabulary issues. Recent work has shown that models operating directly on raw text at the byte or character level can mitigate these limitations. In this paper, we evaluate two token-free models, ByT5 and CANINE, on the task of sarcasm detection in both social media (Twitter) and non-social media (news headlines) domains. We fine-tune and benchmark these models against token-based baselines and state-of-the-art approaches. Our results show that ByT5-small and CANINE outperform token-based counterparts and achieve new state-of-the-art performance, improving accuracy by 0.77% and 0.49% on the News Headlines and Twitter Sarcasm datasets, respectively. These findings underscore the potential of token-free models for robust NLP in noisy and informal domains such as social media. |
| title | Token-free Models for Sarcasm Detection |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2505.01006 |