Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America
Fuente:
arXiv
Saved in:
| Main Authors: | Karmim, Yannis, Pino, Renato, Contreras, Hernan, Lira, Hernan, Cifuentes, Sebastian, Escoffier, Simon, Martí, Luis, Seddah, Djamé, Barrière, Valentin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers
by: Barriere, Valentin, et al.
Published: (2024)
by: Barriere, Valentin, et al.
Published: (2024)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
by: Antoun, Wissam, et al.
Published: (2025)
by: Antoun, Wissam, et al.
Published: (2025)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021)
by: Riabi, Arij, et al.
Published: (2021)
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023)
by: Antoun, Wissam, et al.
Published: (2023)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language
by: Riabi, Arij, et al.
Published: (2023)
by: Riabi, Arij, et al.
Published: (2023)
Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
by: Lopetegui, Javier A., et al.
Published: (2024)
by: Lopetegui, Javier A., et al.
Published: (2024)
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
by: Mouilleron, Virginie, et al.
Published: (2026)
by: Mouilleron, Virginie, et al.
Published: (2026)
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
The Wikidata Query Logs Dataset
by: Walter, Sebastian, et al.
Published: (2026)
by: Walter, Sebastian, et al.
Published: (2026)
Fantastic Biases (What are They) and Where to Find Them
by: Barriere, Valentin
Published: (2024)
by: Barriere, Valentin
Published: (2024)
Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
by: Hüsünbeyi, Z. Melce, et al.
Published: (2026)
Disentangling meaning from language in LLM-based machine translation
by: Lasnier, Théo, et al.
Published: (2026)
by: Lasnier, Théo, et al.
Published: (2026)
Language-Switching Triggers Take a Latent Detour Through Language Models
by: Kulumba, Francis, et al.
Published: (2026)
by: Kulumba, Francis, et al.
Published: (2026)
Supra-Laplacian Encoding for Transformer on Dynamic Graphs
by: Karmim, Yannis, et al.
Published: (2024)
by: Karmim, Yannis, et al.
Published: (2024)
ITEM: Improving Training and Evaluation of Message-Passing based GNNs for top-k recommendation
by: Karmim, Yannis, et al.
Published: (2024)
by: Karmim, Yannis, et al.
Published: (2024)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
by: Antoun, Wissam, et al.
Published: (2024)
by: Antoun, Wissam, et al.
Published: (2024)
Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
by: Hilel, Mohammed, et al.
Published: (2025)
by: Hilel, Mohammed, et al.
Published: (2025)
Solidification in Square Section.
by: Carlos Hernán Salinas Lira
Published: (2001)
by: Carlos Hernán Salinas Lira
Published: (2001)
Solidificación de aleación de aluminio en cavidad cuadrada
by: Carlos Hernán Salinas Lira
Published: (2006)
by: Carlos Hernán Salinas Lira
Published: (2006)
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
by: Calbucura, Nicolas, et al.
Published: (2025)
by: Calbucura, Nicolas, et al.
Published: (2025)
Deep Natural Language Feature Learning for Interpretable Prediction
by: Urrutia, Felipe, et al.
Published: (2023)
by: Urrutia, Felipe, et al.
Published: (2023)
Latinoamérica entrevistada / Hern n Becerra Pino
by: Becerra Pino, Hernán, 1957-
Published: (1957)
by: Becerra Pino, Hernán, 1957-
Published: (1957)
La palabra y la tinta / Hern n Becerra, Pino
by: Becerra Pino, Hernán, 1957-
Published: (1972)
by: Becerra Pino, Hernán, 1957-
Published: (1972)
Gaperon: A Peppered English-French Generative Language Model Suite
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
Mapping the Past: Geographically Linking an Early 20th Century Swedish Encyclopedia with Wikidata
by: Ahlin, Axel, et al.
Published: (2024)
by: Ahlin, Axel, et al.
Published: (2024)
StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos
by: Barriere, Valentin, et al.
Published: (2025)
by: Barriere, Valentin, et al.
Published: (2025)
The Bias of “Therapeutic Illusion”: Do We Have to Curb It?
by: Hernán C. Doval
Published: (2017)
by: Hernán C. Doval
Published: (2017)
Temporal receptive field in dynamic graph learning: A comprehensive analysis
by: Karmim, Yannis, et al.
Published: (2024)
by: Karmim, Yannis, et al.
Published: (2024)
Scholarly Wikidata: Population and Exploration of Conference Data in Wikidata using LLMs
by: Mihindukulasooriya, Nandana, et al.
Published: (2024)
by: Mihindukulasooriya, Nandana, et al.
Published: (2024)
Editorial
by: Édgar Hernán Fuentes-Contreras
Published: (2021)
by: Édgar Hernán Fuentes-Contreras
Published: (2021)
Superar la sostenibilidad urbana: una ruta para América Latina
by: Christian Hernán Contreras-Escandón
Published: (2017)
by: Christian Hernán Contreras-Escandón
Published: (2017)
Uso de métricas em mídias sociais e indicadores de desempenho do site e sua relação com o valor da marca em empresas de cosméticos no Brasil
by: Luis Hernan Contreras Pinochet
Published: (2018)
by: Luis Hernan Contreras Pinochet
Published: (2018)
Facticidad y acción de tutela: presentación preliminar de un estudio empírico de la formulación y efectos de la acción de tutela en el marco colombiano, entre los años 1992-2011
by: Édgar Hernán Fuentes Contreras
Published: (2014)
by: Édgar Hernán Fuentes Contreras
Published: (2014)
The influence of the attributes of “Internet of Things” products on functional and emotional experiences of purchase intention
by: Luis Hernan Contreras Pinochet
Published: (2018)
by: Luis Hernan Contreras Pinochet
Published: (2018)
AVALIAÇÃO DOS CONSUMIDORES DA COMUNIDADE ACADÊMICA DE UMA INSTITUIÇÃO DE ENSINO SUPERIOR PÚBLICA EM RELAÇÃO AS PRÁTICAS DE TI VERDE NAS ORGANIZAÇÕES
by: Luis Hernan Contreras Pinochet
Published: (2015)
by: Luis Hernan Contreras Pinochet
Published: (2015)
Contribuições para a gestão estratégica de instituições de ciência e tecnologia
by: Hernan Edgardo Contreras Alday
Published: (2011)
by: Hernan Edgardo Contreras Alday
Published: (2011)
GENEALOGÍA DE LA ASIMILACIÓN DE LO NORMATIVO: ANÁLISIS DEL ESTUDIO DEL DERECHO EN LOS INICIOS DE LAS UNIVERSIDADES OCCIDENTALES
by: Édgar Hernán Fuentes Contreras
Published: (2017)
by: Édgar Hernán Fuentes Contreras
Published: (2017)
Similar Items
-
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
by: Riabi, Arij, et al.
Published: (2024) -
A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers
by: Barriere, Valentin, et al.
Published: (2024) -
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
by: Antoun, Wissam, et al.
Published: (2025) -
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
by: Riabi, Arij, et al.
Published: (2021) -
From Text to Source: Results in Detecting Large Language Model-Generated Content
by: Antoun, Wissam, et al.
Published: (2023)