Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition

Fuente: Redalyc
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Hiram Calvo
Format: Artículo científico
Sprache:en
Veröffentlicht: Instituto Politécnico Nacional 2014
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1876481778104926208
author Hiram Calvo
author_facet Hiram Calvo
contents Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition Hiram Calvo Andrea Segura-Olivares Alejandro García Computación grams syntactic n similarity measures dependency analysis constituent analysis Paraphrase recognition consists in detecting if an expression restated as another expression contains the same information. Traditionally, for solving this prob- lem, several lexical, syntactic and semantic based tech- niques are used. For measuring word overlapping, most of the works use n-grams; however syntactic n-grams have been scantily explored. We propose using syntac- tic dependency and constituent n-grams combined with common NLP techniques such as stemming, synonym detection, similarity measures , and linear combination and a similarity matrix built in turn from syntactic n- grams. We measure and compare the performance of our system by using the Microsoft Research Paraphrase Corpus. An in-depth research is presented in order to present the strengths and weaknesses of each ap- proach, as well as a common error analysis section. Our main motivation was to determine which syntactic approach had a better performance for this task: syn- tactic dependency n-grams, or syntactic constituent n- grams. We compare too both approaches with traditional n-grams and state-of-the-art systems. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067009 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18
format Artículo científico
id redalyc_61532067009
institution Redalyc
language en
publishDate 2014
publisher Instituto Politécnico Nacional
spellingShingle Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition
Hiram Calvo
Computación
grams
syntactic n
similarity measures
dependency analysis
constituent analysis
Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition Hiram Calvo Andrea Segura-Olivares Alejandro García Computación grams syntactic n similarity measures dependency analysis constituent analysis Paraphrase recognition consists in detecting if an expression restated as another expression contains the same information. Traditionally, for solving this prob- lem, several lexical, syntactic and semantic based tech- niques are used. For measuring word overlapping, most of the works use n-grams; however syntactic n-grams have been scantily explored. We propose using syntac- tic dependency and constituent n-grams combined with common NLP techniques such as stemming, synonym detection, similarity measures , and linear combination and a similarity matrix built in turn from syntactic n- grams. We measure and compare the performance of our system by using the Microsoft Research Paraphrase Corpus. An in-depth research is presented in order to present the strengths and weaknesses of each ap- proach, as well as a common error analysis section. Our main motivation was to determine which syntactic approach had a better performance for this task: syn- tactic dependency n-grams, or syntactic constituent n- grams. We compare too both approaches with traditional n-grams and state-of-the-art systems. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067009 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18
title Dependency vs. Constituent Based Syntactic N-Grams in Text Similarity Measures for Paraphrase Recognition
topic Computación
grams
syntactic n
similarity measures
dependency analysis
constituent analysis
url https://www.redalyc.org/articulo.oa?id=61532067009