Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition
Fuente:
Redalyc
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Artículo científico |
| Sprache: | en |
| Veröffentlicht: |
Instituto Politécnico Nacional
2014
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1876481778104926208 |
|---|---|
| author | Hiram Calvo |
| author_facet | Hiram Calvo |
| contents | Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition Hiram Calvo Andrea Segura-Olivares Alejandro García Computación grams syntactic n similarity measures dependency analysis constituent analysis Paraphrase recognition consists in detecting if an expression restated as another expression contains the same information. Traditionally, for solving this prob- lem, several lexical, syntactic and semantic based tech- niques are used. For measuring word overlapping, most of the works use n-grams; however syntactic n-grams have been scantily explored. We propose using syntac- tic dependency and constituent n-grams combined with common NLP techniques such as stemming, synonym detection, similarity measures , and linear combination and a similarity matrix built in turn from syntactic n- grams. We measure and compare the performance of our system by using the Microsoft Research Paraphrase Corpus. An in-depth research is presented in order to present the strengths and weaknesses of each ap- proach, as well as a common error analysis section. Our main motivation was to determine which syntactic approach had a better performance for this task: syn- tactic dependency n-grams, or syntactic constituent n- grams. We compare too both approaches with traditional n-grams and state-of-the-art systems. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067009 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18 |
| format | Artículo científico |
| id | redalyc_61532067009 |
| institution | Redalyc |
| language | en |
| publishDate | 2014 |
| publisher | Instituto Politécnico Nacional |
| spellingShingle | Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition Hiram Calvo Computación grams syntactic n similarity measures dependency analysis constituent analysis Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition Hiram Calvo Andrea Segura-Olivares Alejandro García Computación grams syntactic n similarity measures dependency analysis constituent analysis Paraphrase recognition consists in detecting if an expression restated as another expression contains the same information. Traditionally, for solving this prob- lem, several lexical, syntactic and semantic based tech- niques are used. For measuring word overlapping, most of the works use n-grams; however syntactic n-grams have been scantily explored. We propose using syntac- tic dependency and constituent n-grams combined with common NLP techniques such as stemming, synonym detection, similarity measures , and linear combination and a similarity matrix built in turn from syntactic n- grams. We measure and compare the performance of our system by using the Microsoft Research Paraphrase Corpus. An in-depth research is presented in order to present the strengths and weaknesses of each ap- proach, as well as a common error analysis section. Our main motivation was to determine which syntactic approach had a better performance for this task: syn- tactic dependency n-grams, or syntactic constituent n- grams. We compare too both approaches with traditional n-grams and state-of-the-art systems. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067009 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18 |
| title | Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition |
| topic | Computación grams syntactic n similarity measures dependency analysis constituent analysis |
| url | https://www.redalyc.org/articulo.oa?id=61532067009 |