Data interference: emojis, homoglyphs, and issues of data fidelity in corpora and their results
Fuente:
arXiv
Saved in:
| Main Author: | Di Cristofaro, Matteo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Impact of emoji exclusion on the performance of Arabic sarcasm detection models
by: Aleryani, Ghalyah H., et al.
Published: (2024)
by: Aleryani, Ghalyah H., et al.
Published: (2024)
Creating emoji lexica from unsupervised sentiment analysis of their descriptions
by: Fernández-Gavilanes, Milagros, et al.
Published: (2024)
by: Fernández-Gavilanes, Milagros, et al.
Published: (2024)
Two CFG Nahuatl for automatic corpora expansion
by: Guzmán-Landa, Juan-José, et al.
Published: (2025)
by: Guzmán-Landa, Juan-José, et al.
Published: (2025)
Multilingual corpora for the study of new concepts in the social sciences and humanities:
by: Kyriakoglou, Revekka, et al.
Published: (2025)
by: Kyriakoglou, Revekka, et al.
Published: (2025)
Curating corpora with classifiers: A case study of clean energy sentiment online
by: Arnold, Michael V., et al.
Published: (2023)
by: Arnold, Michael V., et al.
Published: (2023)
Entropy and type-token ratio in gigaword corpora
by: Rosillo-Rodes, Pablo, et al.
Published: (2024)
by: Rosillo-Rodes, Pablo, et al.
Published: (2024)
DHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling
by: Fedorova, Mariia, et al.
Published: (2026)
by: Fedorova, Mariia, et al.
Published: (2026)
Large corpora and large language models: a replicable method for automating grammatical annotation
by: Morin, Cameron, et al.
Published: (2024)
by: Morin, Cameron, et al.
Published: (2024)
Machine learning and emoji prediction: How much accuracy can MARBERT achieve?
by: Shormani, Mohammed Q., et al.
Published: (2026)
by: Shormani, Mohammed Q., et al.
Published: (2026)
Dynamic Embedded Topic Models: properties and recommendations based on diverse corpora
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
Language corpora for the Dutch medical domain
by: van Es, B.
Published: (2026)
by: van Es, B.
Published: (2026)
Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo
by: Mbogho, Audrey, et al.
Published: (2025)
by: Mbogho, Audrey, et al.
Published: (2025)
Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
by: De Cristofaro, Domenico, et al.
Published: (2025)
by: De Cristofaro, Domenico, et al.
Published: (2025)
Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation
by: Esplà-Gomis, Miquel, et al.
Published: (2024)
by: Esplà-Gomis, Miquel, et al.
Published: (2024)
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
by: De Cristofaro, Domenico, et al.
Published: (2026)
by: De Cristofaro, Domenico, et al.
Published: (2026)
A Systematic Review of Federated Generative Models
by: Gargary, Ashkan Vedadi, et al.
Published: (2024)
by: Gargary, Ashkan Vedadi, et al.
Published: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
by: Gupta, Ankit, et al.
Published: (2024)
by: Gupta, Ankit, et al.
Published: (2024)
Support-verb constructions in the corpora of Greek
by: Fendel, Victoria B.
Published: (2025)
by: Fendel, Victoria B.
Published: (2025)
Identifying economic narratives in large text corpora -- An integrated approach using Large Language Models
by: Schmidt, Tobias, et al.
Published: (2025)
by: Schmidt, Tobias, et al.
Published: (2025)
Exploring digitally-mediated communication with corpora
Published: (2025)
Published: (2025)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
by: Vázquez, Alain, et al.
Published: (2026)
by: Vázquez, Alain, et al.
Published: (2026)
Automatic Logical Forms improve fidelity in Table-to-Text generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
by: Pouplin, Thomas, et al.
Published: (2024)
by: Pouplin, Thomas, et al.
Published: (2024)
Coconstructions in spoken data: UD annotation guidelines and first results
by: Pannitto, Ludovica, et al.
Published: (2026)
by: Pannitto, Ludovica, et al.
Published: (2026)
Explaining word embeddings with perfect fidelity: Case study in research impact prediction
by: Dvorackova, Lucie, et al.
Published: (2024)
by: Dvorackova, Lucie, et al.
Published: (2024)
Dual use issues in the field of Natural Language Generation
by: van Miltenburg, Emiel
Published: (2025)
by: van Miltenburg, Emiel
Published: (2025)
"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs
by: Costa, Mariana Lins
Published: (2025)
by: Costa, Mariana Lins
Published: (2025)
Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training
by: Yang, Yanlai, et al.
Published: (2024)
by: Yang, Yanlai, et al.
Published: (2024)
Applications, challenges and ethical issues of AI and ChatGPT in education
by: Sidiropoulos, Dimitrios, et al.
Published: (2024)
by: Sidiropoulos, Dimitrios, et al.
Published: (2024)
Synthetic Data: Methods, Use Cases, and Risks
by: De Cristofaro, Emiliano
Published: (2023)
by: De Cristofaro, Emiliano
Published: (2023)
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
by: He, Muyu, et al.
Published: (2025)
by: He, Muyu, et al.
Published: (2025)
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
by: Bono, Carlo, et al.
Published: (2025)
by: Bono, Carlo, et al.
Published: (2025)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
by: Lyth, Dan, et al.
Published: (2024)
by: Lyth, Dan, et al.
Published: (2024)
Hey, wait a minute: on at-issue sensitivity in Language Models
by: Kim, Sanghee J., et al.
Published: (2025)
by: Kim, Sanghee J., et al.
Published: (2025)
Preliminary Report: Enhancing Role Differentiation in Conversational HCI Through Chromostereopsis
by: Grella, Matteo
Published: (2025)
by: Grella, Matteo
Published: (2025)
LLMs with Personalities in Multi-issue Negotiation Games
by: Noh, Sean, et al.
Published: (2024)
by: Noh, Sean, et al.
Published: (2024)
Improving Large Language Model (LLM) fidelity through context-aware grounding: A systematic approach to reliability and veracity
by: Talukdar, Wrick, et al.
Published: (2024)
by: Talukdar, Wrick, et al.
Published: (2024)
Lost in the Pipeline: How Well Do Large Language Models Handle Data Preparation?
by: Spreafico, Matteo, et al.
Published: (2025)
by: Spreafico, Matteo, et al.
Published: (2025)
Quantum Attention by Overlap Interference: Predicting Sequences from Classical and Many-Body Quantum Data
by: Pecilli, Alessio, et al.
Published: (2026)
by: Pecilli, Alessio, et al.
Published: (2026)
Similar Items
-
Impact of emoji exclusion on the performance of Arabic sarcasm detection models
by: Aleryani, Ghalyah H., et al.
Published: (2024) -
Creating emoji lexica from unsupervised sentiment analysis of their descriptions
by: Fernández-Gavilanes, Milagros, et al.
Published: (2024) -
Two CFG Nahuatl for automatic corpora expansion
by: Guzmán-Landa, Juan-José, et al.
Published: (2025) -
Multilingual corpora for the study of new concepts in the social sciences and humanities:
by: Kyriakoglou, Revekka, et al.
Published: (2025) -
Curating corpora with classifiers: A case study of clean energy sentiment online
by: Arnold, Michael V., et al.
Published: (2023)