Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Jagadeeshan, Manoj Balaji, Raj, Prince, Goyal, Pawan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIRudit: Cross-Lingual Information Retrieval of Scientific Documents
by: Valentini, Francisco, et al.
Published: (2025)
by: Valentini, Francisco, et al.
Published: (2025)
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
by: Kumar, Sujeet, et al.
Published: (2025)
by: Kumar, Sujeet, et al.
Published: (2025)
Generative Query Expansion with Multilingual LLMs for Cross-Lingual Information Retrieval
by: Macmillan-Scott, Olivia, et al.
Published: (2025)
by: Macmillan-Scott, Olivia, et al.
Published: (2025)
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
Beyond Ranked Lists: The SARAL Framework for Cross-Lingual Document Set Retrieval
by: Agarwal, Shantanu, et al.
Published: (2025)
by: Agarwal, Shantanu, et al.
Published: (2025)
Backtracing: Retrieving the Cause of the Query
by: Wang, Rose E., et al.
Published: (2024)
by: Wang, Rose E., et al.
Published: (2024)
Evaluating Large Language Models for Cross-Lingual Retrieval
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora
by: Amiraz, Chen, et al.
Published: (2025)
by: Amiraz, Chen, et al.
Published: (2025)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
Bridging Language Gaps: Advances in Cross-Lingual Information Retrieval with Multilingual LLMs
by: Goworek, Roksana, et al.
Published: (2025)
by: Goworek, Roksana, et al.
Published: (2025)
Improving Document Retrieval Coherence for Semantically Equivalent Queries
by: Campese, Stefano, et al.
Published: (2025)
by: Campese, Stefano, et al.
Published: (2025)
Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
by: Guo, Hao, et al.
Published: (2025)
by: Guo, Hao, et al.
Published: (2025)
STARD: A Chinese Statute Retrieval Dataset with Real Queries Issued by Non-professionals
by: Su, Weihang, et al.
Published: (2024)
by: Su, Weihang, et al.
Published: (2024)
DAPR: A Benchmark on Document-Aware Passage Retrieval
by: Wang, Kexin, et al.
Published: (2023)
by: Wang, Kexin, et al.
Published: (2023)
On The Persona-based Summarization of Domain-Specific Documents
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training
by: Lee, Seungyoon, et al.
Published: (2026)
by: Lee, Seungyoon, et al.
Published: (2026)
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers
by: Sawarkar, Kunal, et al.
Published: (2024)
by: Sawarkar, Kunal, et al.
Published: (2024)
A Pointer Network-based Approach for Joint Extraction and Detection of Multi-Label Multi-Class Intents
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
QueryBuilder: Human-in-the-Loop Query Development for Information Retrieval
by: Kandula, Hemanth, et al.
Published: (2024)
by: Kandula, Hemanth, et al.
Published: (2024)
Enhancing Retrieval in QA Systems with Derived Feature Association
by: Shah, Keyush, et al.
Published: (2024)
by: Shah, Keyush, et al.
Published: (2024)
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking
by: Vonášek, Josef, et al.
Published: (2024)
by: Vonášek, Josef, et al.
Published: (2024)
GOLFer: Smaller LM-Generated Documents Hallucination Filter & Combiner for Query Expansion in Information Retrieval
by: Liu, Lingyuan, et al.
Published: (2025)
by: Liu, Lingyuan, et al.
Published: (2025)
Building Russian Benchmark for Evaluation of Information Retrieval Models
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
by: Chen, Jianlyu, et al.
Published: (2024)
by: Chen, Jianlyu, et al.
Published: (2024)
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
by: Li, Xiangyang, et al.
Published: (2024)
by: Li, Xiangyang, et al.
Published: (2024)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
by: Pereira, Jayr, et al.
Published: (2026)
by: Pereira, Jayr, et al.
Published: (2026)
Database-Augmented Query Representation for Information Retrieval
by: Jeong, Soyeong, et al.
Published: (2024)
by: Jeong, Soyeong, et al.
Published: (2024)
IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
by: Paul, Shounak, et al.
Published: (2025)
by: Paul, Shounak, et al.
Published: (2025)
Qilin: A Multimodal Information Retrieval Dataset with APP-level User Sessions
by: Chen, Jia, et al.
Published: (2025)
by: Chen, Jia, et al.
Published: (2025)
PosIR: Position-Aware Heterogeneous Information Retrieval Benchmark
by: Zeng, Ziyang, et al.
Published: (2026)
by: Zeng, Ziyang, et al.
Published: (2026)
BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives
by: Sinha, Aarush, et al.
Published: (2025)
by: Sinha, Aarush, et al.
Published: (2025)
Mapping Transformer Leveraged Embeddings for Cross-Lingual Document Representation
by: Tashu, Tsegaye Misikir, et al.
Published: (2024)
by: Tashu, Tsegaye Misikir, et al.
Published: (2024)
Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation
by: Ehsan, Toqeer, et al.
Published: (2025)
by: Ehsan, Toqeer, et al.
Published: (2025)
Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
by: Babakhin, Yauhen, et al.
Published: (2025)
by: Babakhin, Yauhen, et al.
Published: (2025)
Similar Items
-
CLIRudit: Cross-Lingual Information Retrieval of Scientific Documents
by: Valentini, Francisco, et al.
Published: (2025) -
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
by: Kumar, Sujeet, et al.
Published: (2025) -
Generative Query Expansion with Multilingual LLMs for Cross-Lingual Information Retrieval
by: Macmillan-Scott, Olivia, et al.
Published: (2025) -
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026) -
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
by: Nehrdich, Sebastian, et al.
Published: (2026)