Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Chhoun, Sovandara, Po, Pichdara, Ros, Sereiwathna, Cho, Wan-Sup, Khoeurn, Saksonita |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Comparative Study of Language Models for Khmer Retrieval-Augmented Question Answering
by: Ros, Sereiwathna, et al.
Published: (2026)
by: Ros, Sereiwathna, et al.
Published: (2026)
Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs
by: Koutsiaris, Christos
Published: (2026)
by: Koutsiaris, Christos
Published: (2026)
CardioEmbed: Domain-Specialized Text Embeddings for Clinical Cardiology
by: Young, Richard J., et al.
Published: (2025)
by: Young, Richard J., et al.
Published: (2025)
Comparison between the Structures of Word Co-occurrence and Word Similarity Networks for Ill-formed and Well-formed Texts in Taiwan Mandarin
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
Session Context Embedding for Intent Understanding in Product Search
by: Mehrdad, Navid, et al.
Published: (2024)
by: Mehrdad, Navid, et al.
Published: (2024)
Local Hybrid Retrieval-Augmented Document QA
by: Astrino, Paolo
Published: (2025)
by: Astrino, Paolo
Published: (2025)
Benchmarking Google Embeddings 2 against Open-Source Models for Multilingual Dense Retrieval and RAG Systems
by: Cirillo, Stefano, et al.
Published: (2026)
by: Cirillo, Stefano, et al.
Published: (2026)
The Case for Intent-Based Query Rewriting
by: Nicolai, Gianna Lisa, et al.
Published: (2025)
by: Nicolai, Gianna Lisa, et al.
Published: (2025)
Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings
by: Yoo, Junseon
Published: (2026)
by: Yoo, Junseon
Published: (2026)
Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing
by: Mutsaddi, Atharva, et al.
Published: (2025)
by: Mutsaddi, Atharva, et al.
Published: (2025)
LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
by: Pradhan, Anu, et al.
Published: (2025)
by: Pradhan, Anu, et al.
Published: (2025)
Flippi: End To End GenAI Assistant for E-Commerce
by: Rajasekar, Anand A., et al.
Published: (2025)
by: Rajasekar, Anand A., et al.
Published: (2025)
Fine-Grained Emotion Recognition via In-Context Learning
by: Ren, Zhaochun, et al.
Published: (2025)
by: Ren, Zhaochun, et al.
Published: (2025)
Adaptive ToR: Complexity-Aware Tree-Based Retrieval for Pareto-Optimal Multi-Intent NLU
by: Yoo, Hee-Kyong, et al.
Published: (2026)
by: Yoo, Hee-Kyong, et al.
Published: (2026)
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
NewsScope: Schema-Grounded Cross-Domain News Claim Extraction with Open Models
by: Pandya, Nidhi
Published: (2025)
by: Pandya, Nidhi
Published: (2025)
Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables
by: Shen, Chen
Published: (2026)
by: Shen, Chen
Published: (2026)
BridgeRAG: Training-Free Bridge-Conditioned Retrieval for Multi-Hop Question Answering
by: Bacellar, Andre
Published: (2026)
by: Bacellar, Andre
Published: (2026)
Train Once, Use Flexibly: A Modular Framework for Multi-Aspect Neural News Recommendation
by: Iana, Andreea, et al.
Published: (2023)
by: Iana, Andreea, et al.
Published: (2023)
Towards Adaptive Context Management for Intelligent Conversational Question Answering
by: Perera, Manoj Madushanka, et al.
Published: (2025)
by: Perera, Manoj Madushanka, et al.
Published: (2025)
Scaling Multilingual Semantic Search in Uber Eats Delivery
by: Ling, Bo, et al.
Published: (2026)
by: Ling, Bo, et al.
Published: (2026)
When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory
by: Shao, Jiaqi, et al.
Published: (2026)
by: Shao, Jiaqi, et al.
Published: (2026)
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Optimizing Retrieval-Augmented Generation for Electrical Engineering: A Case Study on ABB Circuit Breakers
by: Alawadhi, Salahuddin, et al.
Published: (2025)
by: Alawadhi, Salahuddin, et al.
Published: (2025)
2024 Google Scholar Research Interest Ranking for Top 3260 Computer Science Authors
by: Rasane, Atharva
Published: (2024)
by: Rasane, Atharva
Published: (2024)
QUARK: Robust Retrieval under Non-Faithful Queries via Query-Anchored Aggregation
by: Lyu, Rita Qiuran, et al.
Published: (2026)
by: Lyu, Rita Qiuran, et al.
Published: (2026)
Improving and Evaluating Open Deep Research Agents
by: Allabadi, Doaa, et al.
Published: (2025)
by: Allabadi, Doaa, et al.
Published: (2025)
FACTUM: Mechanistic Detection of Citation Hallucination in Long-Form RAG
by: Dassen, Maxime, et al.
Published: (2026)
by: Dassen, Maxime, et al.
Published: (2026)
Algorithmic Trust and Compliance: Benchmarking Brand Notability for UK iGaming Entities in Generative Search Engines
by: Oruesagasti, Julen
Published: (2026)
by: Oruesagasti, Julen
Published: (2026)
Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation
by: Ye, Hua, et al.
Published: (2026)
by: Ye, Hua, et al.
Published: (2026)
Contextually Aware E-Commerce Product Question Answering using RAG
by: Tangarajan, Praveen, et al.
Published: (2025)
by: Tangarajan, Praveen, et al.
Published: (2025)
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance
by: Eyasir, Abrar, et al.
Published: (2026)
by: Eyasir, Abrar, et al.
Published: (2026)
Graph-GRPO: Dependency-Aware Credit Assignment for Generative E-commerce Search Relevance
by: Che, Jiarui, et al.
Published: (2026)
by: Che, Jiarui, et al.
Published: (2026)
Optimizing open-domain question answering with graph-based retrieval augmented generation
by: Cahoon, Joyce, et al.
Published: (2025)
by: Cahoon, Joyce, et al.
Published: (2025)
Promoting Research Collaboration with Open Data Driven Team Recommendation in Response to Call for Proposals
by: Valluru, Siva Likitha, et al.
Published: (2023)
by: Valluru, Siva Likitha, et al.
Published: (2023)
NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data
by: Ming, Cong, et al.
Published: (2026)
by: Ming, Cong, et al.
Published: (2026)
Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship
by: Pan, Yating, et al.
Published: (2026)
by: Pan, Yating, et al.
Published: (2026)
Architecture Matters More Than Scale: A Comparative Study of Retrieval and Memory Augmentation for Financial QA Under SME Compute Constraints
by: Liu, Jianan, et al.
Published: (2026)
by: Liu, Jianan, et al.
Published: (2026)
Evaluating the Effectiveness of Large Language Models in Automated News Article Summarization
by: Houamegni, Lionel Richy Panlap, et al.
Published: (2025)
by: Houamegni, Lionel Richy Panlap, et al.
Published: (2025)
Similar Items
-
A Comparative Study of Language Models for Khmer Retrieval-Augmented Question Answering
by: Ros, Sereiwathna, et al.
Published: (2026) -
Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs
by: Koutsiaris, Christos
Published: (2026) -
CardioEmbed: Domain-Specialized Text Embeddings for Clinical Cardiology
by: Young, Richard J., et al.
Published: (2025) -
Comparison between the Structures of Word Co-occurrence and Word Similarity Networks for Ill-formed and Well-formed Texts in Taiwan Mandarin
by: Huang, Po-Hsuan, et al.
Published: (2024) -
Session Context Embedding for Intent Understanding in Product Search
by: Mehrdad, Navid, et al.
Published: (2024)