Residualized Similarity for Faithfully Explainable Authorship Verification
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915535687516160 |
|---|---|
| author | Zeng, Peter Alipoormolabashi, Pegah Mun, Jihu Dey, Gourab Soni, Nikita Balasubramanian, Niranjan Rambow, Owen Schwartz, H. |
| author_facet | Zeng, Peter Alipoormolabashi, Pegah Mun, Jihu Dey, Gourab Soni, Nikita Balasubramanian, Niranjan Rambow, Owen Schwartz, H. |
| contents | Responsible use of Authorship Verification (AV) systems not only requires high accuracy but also interpretable solutions. More importantly, for systems to be used to make decisions with real-world consequences requires the model's prediction to be explainable using interpretable features that can be traced to the original texts. Neural methods achieve high accuracies, but their representations lack direct interpretability. Furthermore, LLM predictions cannot be explained faithfully -- if there is an explanation given for a prediction, it doesn't represent the reasoning process behind the model's prediction. In this paper, we introduce Residualized Similarity (RS), a novel method that supplements systems using interpretable features with a neural network to improve their performance while maintaining interpretability. Authorship verification is fundamentally a similarity task, where the goal is to measure how alike two documents are. The key idea is to use the neural network to predict a similarity residual, i.e. the error in the similarity predicted by the interpretable system. Our evaluation across four datasets shows that not only can we match the performance of state-of-the-art authorship verification models, but we can show how and to what degree the final prediction is faithful and interpretable. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_05362 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Residualized Similarity for Faithfully Explainable Authorship Verification Zeng, Peter Alipoormolabashi, Pegah Mun, Jihu Dey, Gourab Soni, Nikita Balasubramanian, Niranjan Rambow, Owen Schwartz, H. Computation and Language Responsible use of Authorship Verification (AV) systems not only requires high accuracy but also interpretable solutions. More importantly, for systems to be used to make decisions with real-world consequences requires the model's prediction to be explainable using interpretable features that can be traced to the original texts. Neural methods achieve high accuracies, but their representations lack direct interpretability. Furthermore, LLM predictions cannot be explained faithfully -- if there is an explanation given for a prediction, it doesn't represent the reasoning process behind the model's prediction. In this paper, we introduce Residualized Similarity (RS), a novel method that supplements systems using interpretable features with a neural network to improve their performance while maintaining interpretability. Authorship verification is fundamentally a similarity task, where the goal is to measure how alike two documents are. The key idea is to use the neural network to predict a similarity residual, i.e. the error in the similarity predicted by the interpretable system. Our evaluation across four datasets shows that not only can we match the performance of state-of-the-art authorship verification models, but we can show how and to what degree the final prediction is faithful and interpretable. |
| title | Residualized Similarity for Faithfully Explainable Authorship Verification |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.05362 |