Curating Stopwords in Marathi: A TF-IDF Approach for Improved Text Analysis and Information Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chavan, Rohan, Patil, Gaurav, Madle, Vishal, Joshi, Raviraj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Topic Modeling in Marathi
von: Shinde, Sanket, et al.
Veröffentlicht: (2025)
von: Shinde, Sanket, et al.
Veröffentlicht: (2025)
MahaSQuAD: Bridging Linguistic Divides in Marathi Question-Answering
von: Ghatage, Ruturaj, et al.
Veröffentlicht: (2024)
von: Ghatage, Ruturaj, et al.
Veröffentlicht: (2024)
L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2025)
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2025)
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2023)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2023)
L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models
von: Kowtal, Nidhi, et al.
Veröffentlicht: (2025)
von: Kowtal, Nidhi, et al.
Veröffentlicht: (2025)
Leveraging Parameter Efficient Training Methods for Low Resource Text Classification: A Case Study in Marathi
von: Deshmukh, Pranita, et al.
Veröffentlicht: (2024)
von: Deshmukh, Pranita, et al.
Veröffentlicht: (2024)
L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in Marathi
von: Deshmukh, Pranita, et al.
Veröffentlicht: (2024)
von: Deshmukh, Pranita, et al.
Veröffentlicht: (2024)
L3Cube-MahaNews: News-based Short Text and Long Document Classification Datasets in Marathi
von: Mittal, Saloni, et al.
Veröffentlicht: (2024)
von: Mittal, Saloni, et al.
Veröffentlicht: (2024)
MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models
von: Jadhav, Suramya, et al.
Veröffentlicht: (2025)
von: Jadhav, Suramya, et al.
Veröffentlicht: (2025)
Long Range Named Entity Recognition for Marathi Documents
von: Deshmukh, Pranita, et al.
Veröffentlicht: (2024)
von: Deshmukh, Pranita, et al.
Veröffentlicht: (2024)
Text Categorization Can Enhance Domain-Agnostic Stopword Extraction
von: Turki, Houcemeddine, et al.
Veröffentlicht: (2024)
von: Turki, Houcemeddine, et al.
Veröffentlicht: (2024)
Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing
von: Mutsaddi, Atharva, et al.
Veröffentlicht: (2025)
von: Mutsaddi, Atharva, et al.
Veröffentlicht: (2025)
Improving the Efficiency of Long Document Classification using Sentence Ranking Approach
von: Kokate, Prathamesh, et al.
Veröffentlicht: (2025)
von: Kokate, Prathamesh, et al.
Veröffentlicht: (2025)
Non-Contextual BERT or FastText? A Comparative Analysis
von: Shanbhag, Abhay, et al.
Veröffentlicht: (2024)
von: Shanbhag, Abhay, et al.
Veröffentlicht: (2024)
A Data Selection Approach for Enhancing Low Resource Machine Translation Using Cross-Lingual Sentence Representations
von: Kowtal, Nidhi, et al.
Veröffentlicht: (2024)
von: Kowtal, Nidhi, et al.
Veröffentlicht: (2024)
IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages
von: Endait, Sharvi, et al.
Veröffentlicht: (2025)
von: Endait, Sharvi, et al.
Veröffentlicht: (2025)
Universal Cross-Lingual Text Classification
von: Savant, Riya, et al.
Veröffentlicht: (2024)
von: Savant, Riya, et al.
Veröffentlicht: (2024)
Chain-of-Translation Prompting (CoTR): A Novel Prompting Technique for Low Resource Languages
von: Deshpande, Tejas, et al.
Veröffentlicht: (2024)
von: Deshpande, Tejas, et al.
Veröffentlicht: (2024)
MIPIAD: Multilingual Indirect Prompt Injection Attack Defense with Qwen -- TF-IDF Hybrid and Meta-Ensemble Learning
von: Muhtadi, Al Muhit, et al.
Veröffentlicht: (2026)
von: Muhtadi, Al Muhit, et al.
Veröffentlicht: (2026)
Towards Building Efficient Sentence BERT Models using Layer Pruning
von: Shelke, Anushka, et al.
Veröffentlicht: (2024)
von: Shelke, Anushka, et al.
Veröffentlicht: (2024)
L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2024)
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2024)
TextGram: Towards a better domain-adaptive pretraining
von: Hiwarkhedkar, Sharayu, et al.
Veröffentlicht: (2024)
von: Hiwarkhedkar, Sharayu, et al.
Veröffentlicht: (2024)
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
von: Rohera, Pritika, et al.
Veröffentlicht: (2025)
von: Rohera, Pritika, et al.
Veröffentlicht: (2025)
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2026)
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2026)
L3Cube-IndicQuest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLMs in Indic Context
von: Rohera, Pritika, et al.
Veröffentlicht: (2024)
von: Rohera, Pritika, et al.
Veröffentlicht: (2024)
L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages
von: Tanksale, Nishant, et al.
Veröffentlicht: (2025)
von: Tanksale, Nishant, et al.
Veröffentlicht: (2025)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
von: Kamath, Anusha, et al.
Veröffentlicht: (2025)
von: Kamath, Anusha, et al.
Veröffentlicht: (2025)
On Importance of Code-Mixed Embeddings for Hate Speech Identification
von: Jagdale, Shruti, et al.
Veröffentlicht: (2024)
von: Jagdale, Shruti, et al.
Veröffentlicht: (2024)
On Limitations of LLM as Annotator for Low Resource Languages
von: Jadhav, Suramya, et al.
Veröffentlicht: (2024)
von: Jadhav, Suramya, et al.
Veröffentlicht: (2024)
Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning
von: Khade, Omkar, et al.
Veröffentlicht: (2024)
von: Khade, Omkar, et al.
Veröffentlicht: (2024)
TextAge: A Curated and Diverse Text Dataset for Age Classification
von: Cheekati, Shravan, et al.
Veröffentlicht: (2024)
von: Cheekati, Shravan, et al.
Veröffentlicht: (2024)
Comparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
Efficient Zero-Shot Long Document Classification by Reducing Context Through Sentence Ranking
von: Kokate, Prathamesh, et al.
Veröffentlicht: (2025)
von: Kokate, Prathamesh, et al.
Veröffentlicht: (2025)
On Importance of Pruning and Distillation for Efficient Low Resource NLP
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2024)
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2024)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
TriNER: A Series of Named Entity Recognition Models For Hindi, Bengali & Marathi
von: Dhamaskar, Mohammed Amaan, et al.
Veröffentlicht: (2025)
von: Dhamaskar, Mohammed Amaan, et al.
Veröffentlicht: (2025)
Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study
von: Paul, Rakesh, et al.
Veröffentlicht: (2025)
von: Paul, Rakesh, et al.
Veröffentlicht: (2025)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Topic Modeling in Marathi
von: Shinde, Sanket, et al.
Veröffentlicht: (2025) -
MahaSQuAD: Bridging Linguistic Divides in Marathi Question-Answering
von: Ghatage, Ruturaj, et al.
Veröffentlicht: (2024) -
L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2025) -
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2023) -
L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models
von: Kowtal, Nidhi, et al.
Veröffentlicht: (2025)