Designing and Implementing a Cross-Language Information Retrieval System Using Linguistic Corpora
Fuente:
Zenodo
Gespeichert in:
| Hauptverfasser: | , |
|---|---|
| Format: | Recurso digital |
| Sprache: | Persisch |
| Veröffentlicht: |
Zenodo
2012
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866902145710686208 |
|---|---|
| author | Nezarat, Amin Mosavi Miangah, Tayebeh |
| author_facet | Nezarat, Amin Mosavi Miangah, Tayebeh |
| contents | Information retrieval (IR) is a crucial area of natural language processing (NLP) and can be defined as finding documents whose content is relevant to the query need of a user. Cross-language information retrieval (CLIR) refers to a kind of information retrieval in which the language of the query and that of searched document are different. In fact, it is a retrieval process where the user presents queries in one language to retrieve documents in another language. This paper tried to construct a bilingual lexicon of parallel chunks of English and Persian from two very large monolingual corpora an English-Persian parallel corpus which could be directly applied to cross-language information retrieval tasks. For this purpose, a statistical measure known as Association Score (AS) was used to compute the association value between every two corresponding chunks in the corpus using a couple of complicated algorithms. Once the CLIR system was developed using this bilingual lexicon, an experiment was performed on a set of one hundred English and Persian phrases and collocations to see to what extend this system was effective in assisting the users find the most relevant and suitable equivalents of their queries in either language. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17057425 |
| institution | Zenodo |
| language | fas |
| publishDate | 2012 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Designing and Implementing a Cross-Language Information Retrieval System Using Linguistic Corpora Nezarat, Amin Mosavi Miangah, Tayebeh Cross-language information retrieval linguistic corpora automated translation intelligent factors Information retrieval (IR) is a crucial area of natural language processing (NLP) and can be defined as finding documents whose content is relevant to the query need of a user. Cross-language information retrieval (CLIR) refers to a kind of information retrieval in which the language of the query and that of searched document are different. In fact, it is a retrieval process where the user presents queries in one language to retrieve documents in another language. This paper tried to construct a bilingual lexicon of parallel chunks of English and Persian from two very large monolingual corpora an English-Persian parallel corpus which could be directly applied to cross-language information retrieval tasks. For this purpose, a statistical measure known as Association Score (AS) was used to compute the association value between every two corresponding chunks in the corpus using a couple of complicated algorithms. Once the CLIR system was developed using this bilingual lexicon, an experiment was performed on a set of one hundred English and Persian phrases and collocations to see to what extend this system was effective in assisting the users find the most relevant and suitable equivalents of their queries in either language. |
| title | Designing and Implementing a Cross-Language Information Retrieval System Using Linguistic Corpora |
| topic | Cross-language information retrieval linguistic corpora automated translation intelligent factors |
| url | https://doi.org/10.5281/zenodo.17057425 |