Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus
Fuente:
arXiv
Guardado en:
| Autores principales: | Messner, Craig, Lippincott, Tom |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pairing Orthographically Variant Literary Words to Standard Equivalents Using Neural Edit Distance Models
por: Messner, Craig, et al.
Publicado: (2024)
por: Messner, Craig, et al.
Publicado: (2024)
Transferring Extreme Subword Style Using Ngram Model-Based Logit Scaling
por: Messner, Craig, et al.
Publicado: (2025)
por: Messner, Craig, et al.
Publicado: (2025)
Pretraining Language Models for Diachronic Linguistic Change Discovery
por: Fittschen, Elisabeth, et al.
Publicado: (2025)
por: Fittschen, Elisabeth, et al.
Publicado: (2025)
Graph-Convolutional Autoencoder Ensembles for the Humanities, Illustrated with a Study of the American Slave Trade
por: Lippincott, Tom
Publicado: (2024)
por: Lippincott, Tom
Publicado: (2024)
Detecting Structured Language Alternations in Historical Documents by Combining Language Identification with Fourier Analysis
por: Sirin, Hale, et al.
Publicado: (2024)
por: Sirin, Hale, et al.
Publicado: (2024)
Dynamic embedded topic models and change-point detection for exploring literary-historical hypotheses
por: Sirin, Hale, et al.
Publicado: (2024)
por: Sirin, Hale, et al.
Publicado: (2024)
Revisiting Common Assumptions about Arabic Dialects in NLP
por: Keleg, Amr, et al.
Publicado: (2025)
por: Keleg, Amr, et al.
Publicado: (2025)
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
por: Dai, Yuhang, et al.
Publicado: (2025)
por: Dai, Yuhang, et al.
Publicado: (2025)
Literary and Colloquial Tamil Dialect Identification
por: Nanmalar, M., et al.
Publicado: (2024)
por: Nanmalar, M., et al.
Publicado: (2024)
Tarab: A Multi-Dialect Corpus of Arabic Lyrics and Poetry
por: El-Haj, Mo
Publicado: (2026)
por: El-Haj, Mo
Publicado: (2026)
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models
por: Bafna, Niyati, et al.
Publicado: (2025)
por: Bafna, Niyati, et al.
Publicado: (2025)
Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces
por: McGovern, Hope, et al.
Publicado: (2025)
por: McGovern, Hope, et al.
Publicado: (2025)
Computational Discovery of Chiasmus in Ancient Religious Text
por: McGovern, Hope, et al.
Publicado: (2025)
por: McGovern, Hope, et al.
Publicado: (2025)
Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
por: Oberkircher, Lena S., et al.
Publicado: (2026)
por: Oberkircher, Lena S., et al.
Publicado: (2026)
A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language Models
por: Park, Namu, et al.
Publicado: (2024)
por: Park, Namu, et al.
Publicado: (2024)
FASSILA: A Corpus for Algerian Dialect Fake News Detection and Sentiment Analysis
por: Abdedaiem, Amin, et al.
Publicado: (2024)
por: Abdedaiem, Amin, et al.
Publicado: (2024)
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus
por: Du, Yexing, et al.
Publicado: (2026)
por: Du, Yexing, et al.
Publicado: (2026)
ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging
por: Nacar, Omer, et al.
Publicado: (2026)
por: Nacar, Omer, et al.
Publicado: (2026)
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
por: Hassan, Md. Rezuwan, et al.
Publicado: (2025)
por: Hassan, Md. Rezuwan, et al.
Publicado: (2025)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
por: McGovern, Hope, et al.
Publicado: (2026)
por: McGovern, Hope, et al.
Publicado: (2026)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
por: Altakrori, Malik H., et al.
Publicado: (2025)
por: Altakrori, Malik H., et al.
Publicado: (2025)
Dynamic Embedded Topic Models: properties and recommendations based on diverse corpora
por: Fittschen, Elisabeth, et al.
Publicado: (2025)
por: Fittschen, Elisabeth, et al.
Publicado: (2025)
The Language of Interoception: Examining Embodiment and Emotion Through a Corpus of Body Part Mentions
por: Wu, Sophie, et al.
Publicado: (2025)
por: Wu, Sophie, et al.
Publicado: (2025)
Corpus Considerations for Annotator Modeling and Scaling
por: Sarumi, Olufunke O., et al.
Publicado: (2024)
por: Sarumi, Olufunke O., et al.
Publicado: (2024)
From Bytes to Biases: Investigating the Cultural Self-Perception of Large Language Models
por: Messner, Wolfgang, et al.
Publicado: (2023)
por: Messner, Wolfgang, et al.
Publicado: (2023)
AlcLaM: Arabic Dialectal Language Model
por: Ahmed, Murtadha, et al.
Publicado: (2024)
por: Ahmed, Murtadha, et al.
Publicado: (2024)
An Annotated Corpus of Arabic Tweets for Hate Speech Analysis
por: Zaghouani, Wajdi, et al.
Publicado: (2025)
por: Zaghouani, Wajdi, et al.
Publicado: (2025)
The InviTE Corpus: Annotating Invectives in Tudor English Texts for Computational Modeling
por: Spliethoff, Sophie, et al.
Publicado: (2025)
por: Spliethoff, Sophie, et al.
Publicado: (2025)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
por: Goldin, Gili, et al.
Publicado: (2024)
por: Goldin, Gili, et al.
Publicado: (2024)
QuranMorph: Morphologically Annotated Quranic Corpus
por: Akra, Diyam, et al.
Publicado: (2025)
por: Akra, Diyam, et al.
Publicado: (2025)
Evaluating Dialect Robustness of Language Models via Conversation Understanding
por: Srirag, Dipankar, et al.
Publicado: (2024)
por: Srirag, Dipankar, et al.
Publicado: (2024)
Large Language Models Discriminate Against Speakers of German Dialects
por: Bui, Minh Duc, et al.
Publicado: (2025)
por: Bui, Minh Duc, et al.
Publicado: (2025)
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
por: Khan, Eeham, et al.
Publicado: (2025)
por: Khan, Eeham, et al.
Publicado: (2025)
ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard Translation
por: Ta, Khoa Anh, et al.
Publicado: (2026)
por: Ta, Khoa Anh, et al.
Publicado: (2026)
Literary Evidence Retrieval via Long-Context Language Models
por: Thai, Katherine, et al.
Publicado: (2025)
por: Thai, Katherine, et al.
Publicado: (2025)
Examining the Utility of Self-disclosure Types for Modeling Annotators of Social Norms
por: Henderson, Kieran, et al.
Publicado: (2025)
por: Henderson, Kieran, et al.
Publicado: (2025)
Annotation Guidelines for Corpus Novelties: Part 1 -- Named Entity Recognition
por: Amalvy, Arthur, et al.
Publicado: (2024)
por: Amalvy, Arthur, et al.
Publicado: (2024)
MentalQA: An Annotated Arabic Corpus for Questions and Answers of Mental Healthcare
por: Alhuzali, Hassan, et al.
Publicado: (2024)
por: Alhuzali, Hassan, et al.
Publicado: (2024)
[Lions: 1] and [Tigers: 2] and [Bears: 3], Oh My! Literary Coreference Annotation with LLMs
por: Hicke, Rebecca M. M., et al.
Publicado: (2024)
por: Hicke, Rebecca M. M., et al.
Publicado: (2024)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
por: Fleisig, Eve, et al.
Publicado: (2024)
por: Fleisig, Eve, et al.
Publicado: (2024)
Ejemplares similares
-
Pairing Orthographically Variant Literary Words to Standard Equivalents Using Neural Edit Distance Models
por: Messner, Craig, et al.
Publicado: (2024) -
Transferring Extreme Subword Style Using Ngram Model-Based Logit Scaling
por: Messner, Craig, et al.
Publicado: (2025) -
Pretraining Language Models for Diachronic Linguistic Change Discovery
por: Fittschen, Elisabeth, et al.
Publicado: (2025) -
Graph-Convolutional Autoencoder Ensembles for the Humanities, Illustrated with a Study of the American Slave Trade
por: Lippincott, Tom
Publicado: (2024) -
Detecting Structured Language Alternations in Historical Documents by Combining Language Identification with Fourier Analysis
por: Sirin, Hale, et al.
Publicado: (2024)