Semantically Cohesive Word Grouping in Indian Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Karthika, N J, Patra, Adyasha, Naidu, Nagasai Saketh, Bhattacharya, Arnab, Ramakrishnan, Ganesh, Dangarikar, Chaitali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
by: J, Karthika N, et al.
Published: (2024)
by: J, Karthika N, et al.
Published: (2024)
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
by: Karthika, N J, et al.
Published: (2025)
by: Karthika, N J, et al.
Published: (2025)
Subjective Behaviors and Preferences in LLM: Language of Browsing
by: Sundaresan, Sai, et al.
Published: (2025)
by: Sundaresan, Sai, et al.
Published: (2025)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2025)
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2025)
MorphTok: Morphologically Grounded Tokenization for Indian Languages
by: Brahma, Maharaj, et al.
Published: (2025)
by: Brahma, Maharaj, et al.
Published: (2025)
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
by: Karthika, N J, et al.
Published: (2026)
by: Karthika, N J, et al.
Published: (2026)
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
LexGen: Domain-aware Multilingual Lexicon Generation
by: Maheshwari, Ayush, et al.
Published: (2024)
by: Maheshwari, Ayush, et al.
Published: (2024)
INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects
by: Sharma, Tarun, et al.
Published: (2026)
by: Sharma, Tarun, et al.
Published: (2026)
Ayn: A Tiny yet Competitive Indian Legal Language Model Pretrained from Scratch
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
Riddle Quest : The Enigma of Words
by: Parasa, Niharika Sri, et al.
Published: (2026)
by: Parasa, Niharika Sri, et al.
Published: (2026)
A2TTS: TTS for Low Resource Indian Languages
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
PARAMANU-GANITA: Can Small Math Language Models Rival with Large Language Models on Mathematical Reasoning?
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training
by: Ponkshe, Kaustubh, et al.
Published: (2024)
by: Ponkshe, Kaustubh, et al.
Published: (2024)
Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation
by: Bhattacharyya, Pramit, et al.
Published: (2025)
by: Bhattacharyya, Pramit, et al.
Published: (2025)
BanglaByT5: Byte-Level Modelling for Bangla
by: Bhattacharyya, Pramit, et al.
Published: (2025)
by: Bhattacharyya, Pramit, et al.
Published: (2025)
BhashaSetu: Cross-Lingual Knowledge Transfer from High-Resource to Extreme Low-Resource Languages
by: Maji, Subhadip, et al.
Published: (2026)
by: Maji, Subhadip, et al.
Published: (2026)
A Likelihood Ratio Test of Genetic Relationship among Languages
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
by: Bhattacharyya, Pramit, et al.
Published: (2024)
by: Bhattacharyya, Pramit, et al.
Published: (2024)
IBPS: Indian Bail Prediction System
by: Srivastava, Puspesh Kumar, et al.
Published: (2025)
by: Srivastava, Puspesh Kumar, et al.
Published: (2025)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022)
by: Maheshwari, Ayush, et al.
Published: (2022)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
by: Thakur, Aamod, et al.
Published: (2025)
by: Thakur, Aamod, et al.
Published: (2025)
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
by: Pandey, Isha, et al.
Published: (2026)
by: Pandey, Isha, et al.
Published: (2026)
TFL: Targeted Bit-Flip Attack on Large Language Model
by: Guo, Jingkai, et al.
Published: (2026)
by: Guo, Jingkai, et al.
Published: (2026)
The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations
by: Soler, Aina Garí, et al.
Published: (2024)
by: Soler, Aina Garí, et al.
Published: (2024)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
by: Guo, Jingkai, et al.
Published: (2025)
by: Guo, Jingkai, et al.
Published: (2025)
Language translation, and change of accent for speech-to-speech task using diffusion model
by: Mishra, Abhishek, et al.
Published: (2025)
by: Mishra, Abhishek, et al.
Published: (2025)
A Code Comprehension Benchmark for Large Language Models for Code
by: Havare, Jayant, et al.
Published: (2025)
by: Havare, Jayant, et al.
Published: (2025)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models
by: Racha, Suraj, et al.
Published: (2025)
by: Racha, Suraj, et al.
Published: (2025)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
by: Maharana, Adyasha, et al.
Published: (2023)
by: Maharana, Adyasha, et al.
Published: (2023)
Legal Judgment Reimagined: PredEx and the Rise of Intelligent AI Interpretation in Indian Courts
by: Nigam, Shubham Kumar, et al.
Published: (2024)
by: Nigam, Shubham Kumar, et al.
Published: (2024)
ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
by: M., Yashwanth, et al.
Published: (2025)
by: M., Yashwanth, et al.
Published: (2025)
GUIDEQ: Framework for Guided Questioning for progressive informational collection and classification
by: Mishra, Priya, et al.
Published: (2024)
by: Mishra, Priya, et al.
Published: (2024)
Evaluating Discourse Cohesion in Pre-trained Language Models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
NyayaAnumana & INLegalLlama: The Largest Indian Legal Judgment Prediction Dataset and Specialized Language Model for Enhanced Decision Analysis
by: Nigam, Shubham Kumar, et al.
Published: (2024)
by: Nigam, Shubham Kumar, et al.
Published: (2024)
Vacaspati: A Diverse Corpus of Bangla Literature
by: Bhattacharyya, Pramit, et al.
Published: (2023)
by: Bhattacharyya, Pramit, et al.
Published: (2023)
Riddle Generation using Learning Resources
by: Parasa, Niharika Sri, et al.
Published: (2023)
by: Parasa, Niharika Sri, et al.
Published: (2023)
REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation
by: Lee, Dong-Ho, et al.
Published: (2025)
by: Lee, Dong-Ho, et al.
Published: (2025)
Similar Items
-
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
by: J, Karthika N, et al.
Published: (2024) -
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
by: Karthika, N J, et al.
Published: (2025) -
Subjective Behaviors and Preferences in LLM: Language of Browsing
by: Sundaresan, Sai, et al.
Published: (2025) -
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2025) -
MorphTok: Morphologically Grounded Tokenization for Indian Languages
by: Brahma, Maharaj, et al.
Published: (2025)