Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
Fuente:
arXiv
Guardado en:
| Autores principales: | Pulipaka, Sidharth, Jain, Sparsh, Sankar, Ashwin, Dabre, Raj |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
por: Sankar, Ashwin, et al.
Publicado: (2024)
por: Sankar, Ashwin, et al.
Publicado: (2024)
PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation
por: Pulipaka, Srikar Kashyap
Publicado: (2026)
por: Pulipaka, Srikar Kashyap
Publicado: (2026)
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
por: Halder, Deepon, et al.
Publicado: (2026)
por: Halder, Deepon, et al.
Publicado: (2026)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
por: Piet, Julien, et al.
Publicado: (2023)
por: Piet, Julien, et al.
Publicado: (2023)
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
por: Ghosh, Poulami, et al.
Publicado: (2024)
por: Ghosh, Poulami, et al.
Publicado: (2024)
Pretraining Language Models Using Translationese
por: Doshi, Meet, et al.
Publicado: (2024)
por: Doshi, Meet, et al.
Publicado: (2024)
When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
por: Elshabrawy, Ahmed, et al.
Publicado: (2025)
por: Elshabrawy, Ahmed, et al.
Publicado: (2025)
Punctuation Prediction for Polish Texts using Transformers
por: Pokrywka, Jakub
Publicado: (2024)
por: Pokrywka, Jakub
Publicado: (2024)
Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP
por: Jayakumar, Thanmay, et al.
Publicado: (2026)
por: Jayakumar, Thanmay, et al.
Publicado: (2026)
IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
por: Khan, Mohammed Safi Ur Rahman, et al.
Publicado: (2024)
por: Khan, Mohammed Safi Ur Rahman, et al.
Publicado: (2024)
PSK@EEUCA 2026: Fine-Tuning Large Language Models with Synthetic Data Augmentation for Multi-Class Toxicity Detection in Gaming Chat
por: Pulipaka, Srikar Kashyap
Publicado: (2026)
por: Pulipaka, Srikar Kashyap
Publicado: (2026)
Resolving Transcription Ambiguity in Spanish: A Hybrid Acoustic-Lexical System for Punctuation Restoration
por: Zhu, Xiliang, et al.
Publicado: (2024)
por: Zhu, Xiliang, et al.
Publicado: (2024)
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
An Empirical Study of In-context Learning in LLMs for Machine Translation
por: Chitale, Pranjal A., et al.
Publicado: (2024)
por: Chitale, Pranjal A., et al.
Publicado: (2024)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
por: Halder, Deepon, et al.
Publicado: (2025)
por: Halder, Deepon, et al.
Publicado: (2025)
A Morphology-Based Investigation of Positional Encodings
por: Ghosh, Poulami, et al.
Publicado: (2024)
por: Ghosh, Poulami, et al.
Publicado: (2024)
Spontaneous Informal Speech Dataset for Punctuation Restoration
por: Liu, Xing Yi, et al.
Publicado: (2024)
por: Liu, Xing Yi, et al.
Publicado: (2024)
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models
por: Atwany, Hanin, et al.
Publicado: (2025)
por: Atwany, Hanin, et al.
Publicado: (2025)
Assessing and Improving Punctuation Robustness in English-Marathi Machine Translation
por: Shejole, Kaustubh Shivshankar, et al.
Publicado: (2025)
por: Shejole, Kaustubh Shivshankar, et al.
Publicado: (2025)
Punctuation and Predicates in Language Models
por: Chauhan, Sonakshi, et al.
Publicado: (2025)
por: Chauhan, Sonakshi, et al.
Publicado: (2025)
Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
por: Sankar, Sanjana, et al.
Publicado: (2025)
por: Sankar, Sanjana, et al.
Publicado: (2025)
Does Dependency Locality Predict Non-canonical Word Order in Hindi?
por: Ranjan, Sidharth, et al.
Publicado: (2024)
por: Ranjan, Sidharth, et al.
Publicado: (2024)
How effective is Multi-source pivoting for Translation of Low Resource Indian Languages?
por: Gaikwad, Pranav, et al.
Publicado: (2024)
por: Gaikwad, Pranav, et al.
Publicado: (2024)
Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification
por: Jeong, Jinhong, et al.
Publicado: (2026)
por: Jeong, Jinhong, et al.
Publicado: (2026)
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
por: Sankar, Ashwin, et al.
Publicado: (2025)
por: Sankar, Ashwin, et al.
Publicado: (2025)
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
por: Koshkin, Roman, et al.
Publicado: (2026)
por: Koshkin, Roman, et al.
Publicado: (2026)
Predicting Punctuation in Ancient Chinese Texts: A Multi-Layered LSTM and Attention-Based Approach
por: Cai, Tracy, et al.
Publicado: (2024)
por: Cai, Tracy, et al.
Publicado: (2024)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
por: Thakur, Aamod, et al.
Publicado: (2025)
por: Thakur, Aamod, et al.
Publicado: (2025)
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
por: Gómez-Zaragozá, Lucía, et al.
Publicado: (2023)
por: Gómez-Zaragozá, Lucía, et al.
Publicado: (2023)
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
por: Prasanjith, Pasunuti, et al.
Publicado: (2025)
por: Prasanjith, Pasunuti, et al.
Publicado: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Punctuation-aware treebank tree binarization
por: Klinger, Eitan, et al.
Publicado: (2025)
por: Klinger, Eitan, et al.
Publicado: (2025)
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems
por: Iakovenko, Olga, et al.
Publicado: (2024)
por: Iakovenko, Olga, et al.
Publicado: (2024)
FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus with Punctuation, Casing, and Context
por: Povey, Anna, et al.
Publicado: (2024)
por: Povey, Anna, et al.
Publicado: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
por: Sharma, Roshan, et al.
Publicado: (2024)
por: Sharma, Roshan, et al.
Publicado: (2024)
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
por: Kaing, Hour, et al.
Publicado: (2025)
por: Kaing, Hour, et al.
Publicado: (2025)
EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech
por: Hassan, Besher, et al.
Publicado: (2026)
por: Hassan, Besher, et al.
Publicado: (2026)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
por: Nareddy, Kartheek Kumar Reddy, et al.
Publicado: (2025)
por: Nareddy, Kartheek Kumar Reddy, et al.
Publicado: (2025)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
por: Seleznyov, Mikhail, et al.
Publicado: (2025)
por: Seleznyov, Mikhail, et al.
Publicado: (2025)
Ejemplares similares
-
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
por: Sankar, Ashwin, et al.
Publicado: (2024) -
PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation
por: Pulipaka, Srikar Kashyap
Publicado: (2026) -
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
por: Saji, Alan, et al.
Publicado: (2025) -
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
por: Halder, Deepon, et al.
Publicado: (2026) -
Mark My Words: Analyzing and Evaluating Language Model Watermarks
por: Piet, Julien, et al.
Publicado: (2023)