Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhattacharyya, Pramit, Bhattacharya, Arnab |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Vacaspati: A Diverse Corpus of Bangla Literature
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2023)
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2023)
BanglaByT5: Byte-Level Modelling for Bangla
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2025)
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2025)
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2024)
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2024)
INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects
di: Sharma, Tarun, et al.
Pubblicazione: (2026)
di: Sharma, Tarun, et al.
Pubblicazione: (2026)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
di: Akavarapu, V. S. D. S. Mahesh, et al.
Pubblicazione: (2025)
di: Akavarapu, V. S. D. S. Mahesh, et al.
Pubblicazione: (2025)
Glitter: Visualizing Lexical Surprisal for Readability in Administrative Texts
di: Černý, Jan, et al.
Pubblicazione: (2026)
di: Černý, Jan, et al.
Pubblicazione: (2026)
BanglaSTEM: A Parallel Corpus for Technical Domain Bangla-English Translation
di: Hasan, Kazi Reyazul, et al.
Pubblicazione: (2025)
di: Hasan, Kazi Reyazul, et al.
Pubblicazione: (2025)
ANUBHUTI: A Comprehensive Corpus For Sentiment Analysis In Bangla Regional Languages
di: Kundu, Swastika, et al.
Pubblicazione: (2025)
di: Kundu, Swastika, et al.
Pubblicazione: (2025)
DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context
di: Sahoo, Pramit, et al.
Pubblicazione: (2025)
di: Sahoo, Pramit, et al.
Pubblicazione: (2025)
Crowdsourcing Lexical Diversity
di: Khalilia, Hadi, et al.
Pubblicazione: (2024)
di: Khalilia, Hadi, et al.
Pubblicazione: (2024)
Abstractive Text Summarization for Bangla Language Using NLP and Machine Learning Approaches
di: Miazee, Asif Ahammad, et al.
Pubblicazione: (2025)
di: Miazee, Asif Ahammad, et al.
Pubblicazione: (2025)
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment
di: Elmadani, Khalid N., et al.
Pubblicazione: (2025)
di: Elmadani, Khalid N., et al.
Pubblicazione: (2025)
Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
di: Arpa, Zaara Zabeen, et al.
Pubblicazione: (2025)
di: Arpa, Zaara Zabeen, et al.
Pubblicazione: (2025)
MONOVAB : An Annotated Corpus for Bangla Multi-label Emotion Detection
di: Banshal, Sumit Kumar, et al.
Pubblicazione: (2023)
di: Banshal, Sumit Kumar, et al.
Pubblicazione: (2023)
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
di: Alam, Kazi Samin Yasar, et al.
Pubblicazione: (2026)
di: Alam, Kazi Samin Yasar, et al.
Pubblicazione: (2026)
LGR2: Language Guided Reward Relabeling for Accelerating Hierarchical Reinforcement Learning
di: Singh, Utsav, et al.
Pubblicazione: (2024)
di: Singh, Utsav, et al.
Pubblicazione: (2024)
Statistical Analysis of Sentence Structures through ASCII, Lexical Alignment and PCA
di: Sahdev, Abhijeet
Pubblicazione: (2025)
di: Sahdev, Abhijeet
Pubblicazione: (2025)
NLP needs Diversity outside of 'Diversity'
di: Tint, Joshua
Pubblicazione: (2026)
di: Tint, Joshua
Pubblicazione: (2026)
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
di: Deshpande, Vijeta, et al.
Pubblicazione: (2025)
di: Deshpande, Vijeta, et al.
Pubblicazione: (2025)
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
di: Nguyen, Thanh-Nhi, et al.
Pubblicazione: (2024)
di: Nguyen, Thanh-Nhi, et al.
Pubblicazione: (2024)
Analyzing Feedback Mechanisms in AI-Generated MCQs: Insights into Readability, Lexical Properties, and Levels of Challenge
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
What is "Typological Diversity" in NLP?
di: Ploeger, Esther, et al.
Pubblicazione: (2024)
di: Ploeger, Esther, et al.
Pubblicazione: (2024)
TartuNLP @ AXOLOTL-24: Leveraging Classifier Output for New Sense Detection in Lexical Semantics
di: Dorkin, Aleksei, et al.
Pubblicazione: (2024)
di: Dorkin, Aleksei, et al.
Pubblicazione: (2024)
The Use of Readability Metrics in Legal Text: A Systematic Literature Review
di: Han, Yu, et al.
Pubblicazione: (2024)
di: Han, Yu, et al.
Pubblicazione: (2024)
Authorship Attribution in Bangla Literature (AABL) via Transfer Learning using ULMFiT
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification
di: Mohsin, Ayesha Afroza, et al.
Pubblicazione: (2025)
di: Mohsin, Ayesha Afroza, et al.
Pubblicazione: (2025)
Accelerating Bangla NLP Tasks with Automatic Mixed Precision: Resource-Efficient Training Preserving Model Efficacy
di: Opi, Md Mehrab Hossain, et al.
Pubblicazione: (2025)
di: Opi, Md Mehrab Hossain, et al.
Pubblicazione: (2025)
TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
Beware of Words: Evaluating the Lexical Diversity of Conversational LLMs using ChatGPT as Case Study
di: Martínez, Gonzalo, et al.
Pubblicazione: (2024)
di: Martínez, Gonzalo, et al.
Pubblicazione: (2024)
Enhancing Task-Oriented Dialogues with Chitchat: a Comparative Study Based on Lexical Diversity and Divergence
di: Stricker, Armand, et al.
Pubblicazione: (2023)
di: Stricker, Armand, et al.
Pubblicazione: (2023)
Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation
di: Ploeger, Esther, et al.
Pubblicazione: (2024)
di: Ploeger, Esther, et al.
Pubblicazione: (2024)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2024)
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2024)
CAIRNS: Balancing Readability and Scientific Accuracy in Climate Adaptation Question Answering
di: Kong, Liangji, et al.
Pubblicazione: (2025)
di: Kong, Liangji, et al.
Pubblicazione: (2025)
BanglaMedQA and BanglaMMedBench: Evaluating Retrieval-Augmented Generation Strategies for Bangla Biomedical Question Answering
di: Sultana, Sadia, et al.
Pubblicazione: (2025)
di: Sultana, Sadia, et al.
Pubblicazione: (2025)
Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
di: Cavalin, Paulo, et al.
Pubblicazione: (2024)
di: Cavalin, Paulo, et al.
Pubblicazione: (2024)
Designing NLP Systems That Adapt to Diverse Worldviews
di: Creanga, Claudiu, et al.
Pubblicazione: (2024)
di: Creanga, Claudiu, et al.
Pubblicazione: (2024)
LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation
di: Go, Gregory Hok Tjoan, et al.
Pubblicazione: (2025)
di: Go, Gregory Hok Tjoan, et al.
Pubblicazione: (2025)
NESTLE: a No-Code Tool for Statistical Analysis of Legal Corpus
di: Cho, Kyoungyeon, et al.
Pubblicazione: (2023)
di: Cho, Kyoungyeon, et al.
Pubblicazione: (2023)
HLDC: Hindi Legal Documents Corpus
di: Kapoor, Arnav, et al.
Pubblicazione: (2022)
di: Kapoor, Arnav, et al.
Pubblicazione: (2022)
A Quantitative Discourse Analysis of Asian Workers in the US Historical Newspapers
di: Park, Jaihyun, et al.
Pubblicazione: (2024)
di: Park, Jaihyun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Vacaspati: A Diverse Corpus of Bangla Literature
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2023) -
BanglaByT5: Byte-Level Modelling for Bangla
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2025) -
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2024) -
INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects
di: Sharma, Tarun, et al.
Pubblicazione: (2026) -
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
di: Akavarapu, V. S. D. S. Mahesh, et al.
Pubblicazione: (2025)