Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Arpa, Zaara Zabeen, Apurbo, Sadnam Sakib, Oishee, Nazia Karim Khan, Abrar, Ajwad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication
by: Ahmad, Arif, et al.
Published: (2024)
by: Ahmad, Arif, et al.
Published: (2024)
Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
by: Abrar, Ajwad, et al.
Published: (2025)
by: Abrar, Ajwad, et al.
Published: (2025)
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
by: Alam, Kazi Samin Yasar, et al.
Published: (2026)
by: Alam, Kazi Samin Yasar, et al.
Published: (2026)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
by: Rafid, Ahmed, et al.
Published: (2026)
by: Rafid, Ahmed, et al.
Published: (2026)
From Chat to Checkup: Can Large Language Models Assist in Diabetes Prediction?
by: Sakib, Shadman, et al.
Published: (2025)
by: Sakib, Shadman, et al.
Published: (2025)
BanglaMedQA and BanglaMMedBench: Evaluating Retrieval-Augmented Generation Strategies for Bangla Biomedical Question Answering
by: Sultana, Sadia, et al.
Published: (2025)
by: Sultana, Sadia, et al.
Published: (2025)
Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach
by: Sani, Ahmed Alfey, et al.
Published: (2026)
by: Sani, Ahmed Alfey, et al.
Published: (2026)
Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies
by: Abrar, Ajwad, et al.
Published: (2025)
by: Abrar, Ajwad, et al.
Published: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
Investigating Transcription Normalization in the Faetar ASR Benchmark
by: Peckham, Leo, et al.
Published: (2025)
by: Peckham, Leo, et al.
Published: (2025)
Vacaspati: A Diverse Corpus of Bangla Literature
by: Bhattacharyya, Pramit, et al.
Published: (2023)
by: Bhattacharyya, Pramit, et al.
Published: (2023)
NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance
by: Eyasir, Abrar, et al.
Published: (2026)
by: Eyasir, Abrar, et al.
Published: (2026)
BanglaSTEM: A Parallel Corpus for Technical Domain Bangla-English Translation
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
by: Semenov, Kirill, et al.
Published: (2025)
by: Semenov, Kirill, et al.
Published: (2025)
Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
by: Altinok, Duygu
Published: (2025)
by: Altinok, Duygu
Published: (2025)
Boosting Disfluency Detection with Large Language Model as Disfluency Generator
by: Cheng, Zhenrong, et al.
Published: (2024)
by: Cheng, Zhenrong, et al.
Published: (2024)
ANUBHUTI: A Comprehensive Corpus For Sentiment Analysis In Bangla Regional Languages
by: Kundu, Swastika, et al.
Published: (2025)
by: Kundu, Swastika, et al.
Published: (2025)
Introducing A Bangla Sentence - Gloss Pair Dataset for Bangla Sign Language Translation and Research
by: Saha, Neelavro, et al.
Published: (2025)
by: Saha, Neelavro, et al.
Published: (2025)
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
by: Kumar, Sujeet, et al.
Published: (2025)
by: Kumar, Sujeet, et al.
Published: (2025)
FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings
by: Muna, Saiyma Sittul, et al.
Published: (2025)
by: Muna, Saiyma Sittul, et al.
Published: (2025)
Social media polarization during conflict: Insights from an ideological stance dataset on Israel-Palestine Reddit comments
by: Ali, Hasin Jawad, et al.
Published: (2025)
by: Ali, Hasin Jawad, et al.
Published: (2025)
MONOVAB : An Annotated Corpus for Bangla Multi-label Emotion Detection
by: Banshal, Sumit Kumar, et al.
Published: (2023)
by: Banshal, Sumit Kumar, et al.
Published: (2023)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs
by: Abrar, Ajwad, et al.
Published: (2025)
by: Abrar, Ajwad, et al.
Published: (2025)
BanglaIPA: Towards Robust Text-to-IPA Transcription with Contextual Rewriting in Bengali
by: Hasan, Jakir, et al.
Published: (2026)
by: Hasan, Jakir, et al.
Published: (2026)
BanglaLlama: LLaMA for Bangla Language
by: Zehady, Abdullah Khan, et al.
Published: (2024)
by: Zehady, Abdullah Khan, et al.
Published: (2024)
Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation
by: Bhattacharyya, Pramit, et al.
Published: (2025)
by: Bhattacharyya, Pramit, et al.
Published: (2025)
CogniAlign: Survivability-Grounded Multi-Agent Moral Reasoning for Safe and Transparent AI
by: Ali, Hasin Jawad, et al.
Published: (2025)
by: Ali, Hasin Jawad, et al.
Published: (2025)
Disfluencies We Live with in Japanese
Published: (2026)
Published: (2026)
Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
by: Simonotti, Martina, et al.
Published: (2026)
by: Simonotti, Martina, et al.
Published: (2026)
Beyond Transcription: Mechanistic Interpretability in ASR
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
by: Lia, Nusrat Jahan, et al.
Published: (2025)
by: Lia, Nusrat Jahan, et al.
Published: (2025)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
Vashantor: A Large-scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language
by: Faria, Fatema Tuj Johora, et al.
Published: (2023)
by: Faria, Fatema Tuj Johora, et al.
Published: (2023)
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Bangla MedER: Multi-BERT Ensemble Approach for the Recognition of Bangla Medical Entity
by: Aurpa, Tanjim Taharat, et al.
Published: (2025)
by: Aurpa, Tanjim Taharat, et al.
Published: (2025)
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
BanglaByT5: Byte-Level Modelling for Bangla
by: Bhattacharyya, Pramit, et al.
Published: (2025)
by: Bhattacharyya, Pramit, et al.
Published: (2025)
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
by: Aggazzotti, Cristina, et al.
Published: (2023)
by: Aggazzotti, Cristina, et al.
Published: (2023)
Similar Items
-
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication
by: Ahmad, Arif, et al.
Published: (2024) -
Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
by: Abrar, Ajwad, et al.
Published: (2025) -
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
by: Alam, Kazi Samin Yasar, et al.
Published: (2026) -
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
by: Rafid, Ahmed, et al.
Published: (2026) -
From Chat to Checkup: Can Large Language Models Assist in Diabetes Prediction?
by: Sakib, Shadman, et al.
Published: (2025)