Paraphrase and Textual Entailment Generation in Czech
Fuente:
Redalyc
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Artículo científico |
| Sprache: | es |
| Veröffentlicht: |
Instituto Politécnico Nacional
2014
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1876431174996328448 |
|---|---|
| author | Zuzana Neverilová |
| author_facet | Zuzana Neverilová |
| contents | Paraphrase and Textual Entailment Generation in Czech Zuzana Neverilová Computación paraphrase textual entailment Games with a purpose natural language generation Paraphrase and textual entailment generation can support natural language processing (NLP) tasks that simulate text understanding, e.g., text summariza- tion, plagiarism detection, or question answering. A paraphrase, i.e., a sentence with the same meaning, conveys a certain piece of information with new words and new syntactic structures. Textual entailment, i.e., an inference that humans will judge most likely true, can employ real-world knowledge in order to make some implicit information explicit. Paraphrases can also be seen as mutual entailments. We present a new system that generates paraphrases and textual entailments from a given text in the Czech language. First, the process is rule-based, i.e., the system analyzes the input text, pro- duces its inner representation, transforms it according to transformation rules, and generates new sentences. Second, the generated sentences are ranked according to a statistical model and only the best ones are output. The decision whether a paraphrase or textual entailment is correct or not is left to humans. For this purpose we designed an annotation game based on a conversation between a detective (the human player) and his assis- tant (the system). The result of such annotation is a collection of annotated pairs text–hypothesis. Currently, the system and the game are intended to collect data in the Czech language. However, the idea can be applied for other languages. So far, we have collected 3,321 H–T pairs. From these pairs, 1,563 were judged cor- rect (47.06 %), 1,238 (37.28 %) were judged incorrect entailments, and 520 (15.66 %) were judged non-sense or unknown. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067010 es http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18 |
| format | Artículo científico |
| id | redalyc_61532067010 |
| institution | Redalyc |
| language | es |
| publishDate | 2014 |
| publisher | Instituto Politécnico Nacional |
| spellingShingle | Paraphrase and Textual Entailment Generation in Czech Zuzana Neverilová Computación paraphrase textual entailment Games with a purpose natural language generation Paraphrase and Textual Entailment Generation in Czech Zuzana Neverilová Computación paraphrase textual entailment Games with a purpose natural language generation Paraphrase and textual entailment generation can support natural language processing (NLP) tasks that simulate text understanding, e.g., text summariza- tion, plagiarism detection, or question answering. A paraphrase, i.e., a sentence with the same meaning, conveys a certain piece of information with new words and new syntactic structures. Textual entailment, i.e., an inference that humans will judge most likely true, can employ real-world knowledge in order to make some implicit information explicit. Paraphrases can also be seen as mutual entailments. We present a new system that generates paraphrases and textual entailments from a given text in the Czech language. First, the process is rule-based, i.e., the system analyzes the input text, pro- duces its inner representation, transforms it according to transformation rules, and generates new sentences. Second, the generated sentences are ranked according to a statistical model and only the best ones are output. The decision whether a paraphrase or textual entailment is correct or not is left to humans. For this purpose we designed an annotation game based on a conversation between a detective (the human player) and his assis- tant (the system). The result of such annotation is a collection of annotated pairs text–hypothesis. Currently, the system and the game are intended to collect data in the Czech language. However, the idea can be applied for other languages. So far, we have collected 3,321 H–T pairs. From these pairs, 1,563 were judged cor- rect (47.06 %), 1,238 (37.28 %) were judged incorrect entailments, and 520 (15.66 %) were judged non-sense or unknown. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067010 es http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18 |
| title | Paraphrase and Textual Entailment Generation in Czech |
| topic | Computación paraphrase textual entailment Games with a purpose natural language generation |
| url | https://www.redalyc.org/articulo.oa?id=61532067010 |