Designing and Implementing a Cross-Language Information Retrieval System Using Linguistic Corpora

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nezarat, Amin, Mosavi Miangah, Tayebeh
Format: Recurso digital
Sprache:Persisch
Veröffentlicht: Zenodo 2012
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902145710686208
author Nezarat, Amin
Mosavi Miangah, Tayebeh
author_facet Nezarat, Amin
Mosavi Miangah, Tayebeh
contents Information retrieval (IR) is a crucial area of natural language processing (NLP) and can be defined as finding documents whose content is relevant to the query need of a user. Cross-language information retrieval (CLIR) refers to a kind of information retrieval in which the language of the query and that of searched document are different. In fact, it is a retrieval process where the user presents queries in one language to retrieve documents in another language. This paper tried to construct a bilingual lexicon of parallel chunks of English and Persian from two very large monolingual corpora an English-Persian parallel corpus which could be directly applied to cross-language information retrieval tasks. For this purpose, a statistical measure known as Association Score (AS) was used to compute the association value between every two corresponding chunks in the corpus using a couple of complicated algorithms. Once the CLIR system was developed using this bilingual lexicon, an experiment was performed on a set of one hundred English and Persian phrases and collocations to see to what extend this system was effective in assisting the users find the most relevant and suitable equivalents of their queries in either language.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17057425
institution Zenodo
language fas
publishDate 2012
publisher Zenodo
record_format zenodo
spellingShingle Designing and Implementing a Cross-Language Information Retrieval System Using Linguistic Corpora
Nezarat, Amin
Mosavi Miangah, Tayebeh
Cross-language information retrieval
linguistic corpora
automated translation
intelligent factors
Information retrieval (IR) is a crucial area of natural language processing (NLP) and can be defined as finding documents whose content is relevant to the query need of a user. Cross-language information retrieval (CLIR) refers to a kind of information retrieval in which the language of the query and that of searched document are different. In fact, it is a retrieval process where the user presents queries in one language to retrieve documents in another language. This paper tried to construct a bilingual lexicon of parallel chunks of English and Persian from two very large monolingual corpora an English-Persian parallel corpus which could be directly applied to cross-language information retrieval tasks. For this purpose, a statistical measure known as Association Score (AS) was used to compute the association value between every two corresponding chunks in the corpus using a couple of complicated algorithms. Once the CLIR system was developed using this bilingual lexicon, an experiment was performed on a set of one hundred English and Persian phrases and collocations to see to what extend this system was effective in assisting the users find the most relevant and suitable equivalents of their queries in either language.
title Designing and Implementing a Cross-Language Information Retrieval System Using Linguistic Corpora
topic Cross-language information retrieval
linguistic corpora
automated translation
intelligent factors
url https://doi.org/10.5281/zenodo.17057425